Friday, April 27, 2018

Setting Up Data Science Projects

At a high level, a data science project has six components in three classes:
  1. Objective 
    1. What is the intended benefit?
    2. How is the problem going to be solved?
  2. Data
    1.  What is the data will be output?
    2. What is the data that will be coming in?
  3. Technical
    1. What will be the analytic approach?
    2. What code will be written?
These aren't steps. Projects often go back and forth between the various components. For instance, we'll realize that the code resources we have available won't make the solution we want possible, so we go all the way back to component 1.2 to see if there is a different way of solving the problem.

The components are ordered in terms of importance! Getting the right problem to solve is vastly more important that writing the best code.

However, almost all of the papers and talks and blogs concentrate on components 3.2. The least important step actually gets the most attention.

That's the problem. I don't know of a good way of getting better at the higher-value stages except to try to always understand what you are doing. For instance, radically changing component 3.1 may mean you are really changing the business problem being addressed; just make sure that is what you want to do.

Friday, April 20, 2018

Lehman's Law

JT Lehman (JT Lehman) has a great rule for setting up data science projects:

If we only knew _____, then we could do _____ and have impact ______. Fill in the blanks.

A lot of project problems take care of themselves if you've got the problem definition right.

Sunday, April 15, 2018

Estimating the 2018 House Elections

Nate Silver (http://fivethirtyeight.com/) made a comment about the 2018 U.S. House of Representative Elections, in that he was expecting modest gains but the long tail for the GOP was really bad. I want to investigate that idea. We don't have enough data to make a hard calculation, but what we can do is to put some solidity around our intuitions.

I'm making a Monte Carlo simulation of the election. The natural way to think about the elections is i comparison with the 2016 Presidential election, so I'm going to start with the Clinton vote in each Congressional district. I'm using CLinton and not Trump because it is easier to think about things from the Democratic point of view.

In the recent special elections, the Democrats have been beating the 2016 Presidential vote by about 17 points on average. The Democrats have been having an ~8 point lead lead over the GOP in the generic House tracker (https://projects.fivethirtyeight.com/congress-generic-ballot-polls) so that leaves a bunch of gap to be explained.

Anecdotally, the Democrats have a marked enthusiasm gap. I have a vague memory of 4 percent, so let's go with that. Also, it seems to me that in the recent special elections the Democrats have been fielding pretty good, above average candidates while the GOP has been fielding bad-to-terrible candidates. This makes sense to me; if I was an ambitious Democrat I would be looking to find a way to get into the game, whereas if I were an ambition Republican I would be finding excuses to sit this one out. So let's give the Democrats a 4 point 'better candidates' boost. This won't apply to races where the GOP incumbent is staying in the race.

This gives us as a baseline matching the recent special elections

Democratic Advantages

Higher Approval Rating                   8 points
Higher Enthusiasm                        4 points
Better Candidates                        4 points
Total                                   16 points

which roughly matches the 17 points we are seeing from the special elections. Why am I breaking the 16 points down like this? Having three buckets makes it a lot easier to think about than having one big lump.

Other Effects

I'm adding a 6-point incumbent advantage. The quantity here is fairly arbitrary; I'd like a good way of getting a better handle on this number in this context. 

I also want to add uncertainty factors. I want one factor to represent uncertainty due to the national environment changing in the next seven months. and another uncertainty factor for race-by-race factors. I'm treating both as normal effects with a mean of 0 and a standard deviation of 2.75. Why 2.75 (which is admittedly a weird number to use)? I'm figuring I want the Democrats to have a 99% chance of getting control of the house in the situation where the general election acts like the special elections, and calibrating the uncertainty to a standard deviation of 2.75 does that.

Incumbent Bonus                          6 points
National Uncertainty                     2.75 s.d.
By-District Uncertainty                  2.75 s.d.

Mainline Results (2018 general matches special elections)

194 is the current Dem seats in the House; 218 is how many seats the Dems need to take control of the House.
so we are looking at about a 60-seat gain on average, and the long tail is pretty brutal for the GOP: a 100-seat loss is quite possible.

Other Options

An advantage of having a model like this is we can change the assumptions and see what the effect is.

Let's start by cutting the basic advantage in half.
Still pretty good, still about an 80% chance of winning the House.. Now let's cut the enthusiasm and candidate bonus in half as well.
Not so good -- only a 40% chance to taking the House.
Let's try moving the base up to 8%, but keeping the enthusiasm and candidate factors at 2.
This is looking pretty good -- about at 88% chance of gaining the House.
This is telling us that it is all about keeping that base advantage; better candidates do not matter that much. 
Lastly, let's try to take away all the Dem advantages:
We let a House that looks a lot like what we got in 2016. This is a decent sanity check for the method.

The Actual Code


# coding: utf-8

# In[1]:


###################################################################
# The goal of this program is to get an idea of the possibilities #
# and spreads of the 2018 house election. We don't have enought   #
# data to make a real prediction, but we can make something that  #
# can give us an idea of the possibilities.                       #
# Recent special elections have been running in the Democrats'    #
# favor, typically doing 15-20 basis points better that the 2016  #
# Trump win. To get a better handle on the wins, we break the Dem #
# Advantage down into three chunks                                #
#     1) Basic favorability advantage; right now that is running  #
#        7-8 basis points in the Dem's favor                      #
#     2) Enthusiasm gap: It seeems like Dems are getting to the   #
#        polls in much higher numbers; I have heard 4 basis points#
#     3) Better candidates. It seems like the Dems have been      #
#        fielding much better candidates that the GOP in the      #
#        recent races; taking 4 basis points here makes the total #
#        Dem advantage match their special election performance   #
# The 'better candidates' factor applies to campaigns where either#
# the Dem is the incumbent, or the GOP incument is not running for#
# some reason.                                                    #
# We also have a factor for incumbency. Right now it it set at 6  #
# basis points; this is out of thin air and a good way to improve #
# the model is to get a better understanding of this factor.      #
# Lastly, we have to factors that represent uncertainty. One      #
# factor is a normal-distributed random number that applies to all#
# the races, representing changes in the national mood; the other #
# is a race-by-race random normal number. Right now both have mean#
# 0 and the same standard deviation. The spread was calibrated to #
# give the Democrats a 99% chance of winning the house under      #
# the most optimistic scenario I considered, which was            #
#     basic +8                                                    #
#     enthusiasm +4                                               #
#     candidate +4                                                #
# This corresponds to the results in the special elections so it  #
# is possible the GOP could do much worse.                        #
###################################################################

import pandas as pd
import numpy as np
import matplotlib as mp
import matplotlib.mlab as mlab
import matplotlib.pyplot as plt


# In[2]:


election=pd.read_csv("C:/work/election2018/election2018v2.txt",sep='\t')


# In[3]:


election.head()


# In[4]:


# This is where the paramters for the simulation get set
# spread = 3.25 goes with incumbent=4.0
# spread = 2.25 goes with incumbent=8.0
# spread = 2.75 goes with incumbent=6.0

demBoost = 8.0
demEnthus = 4.0
candidateFactor=4.0
incumbent=6.0
spread = 2.75
sims = 10000

def demParty(x):
    if x=='Democratic Party':
        return 1
    else:
        return 0

def winf(x):
    if x>50.0:
        return 1
    else:
        return 0
    
winList=[]


# In[5]:


# Running the election simulations
# candidateFactor does not apply if the the incumbent is a Republican who is running.
# We are baing the results off of the Clinton percent in the 2016 election, so the incumbent
# effect for GOP candidates is negative.

for j in range(sims):
    natRand = np.random.normal(0,spread,1)
    disRand = np.random.normal(0,spread,election.shape[0])
    election['demBoost']=demBoost
    election['demEnthus']=demEnthus
    election['candidateFactor'] = ( election['Party'].apply(demParty)*candidateFactor+
                                    (1-election['Party'].apply(demParty))*election['Retiring']*candidateFactor)
    election['incumbent'] = (1-election['Retiring'])*incumbent*(-1+2*election['Party'].apply(demParty))
    election['result']=(election['Clinton']+election['demBoost']+election['demEnthus']+election['candidateFactor']+
                        election['incumbent']+natRand+disRand)
    election['wins'] = election['result'].apply(winf)
    winList.append(election['wins'].sum())
    


# In[6]:


# Create the histograms of the simulations
font = {'weight' : 'bold',
        'size'   : 120}
mp.rc('font', **font)

plt.figure(figsize=(120,120))
plt.hist(winList, 20, density=False,facecolor='blue')
plt.xlabel('Democratic Wins; Green Line is 218, Red Line is 194',fontsize=200)
plt.ylabel('Simulations Out Of '+str(sims),fontsize=200)
title1='Democratic House Wins: Base Advantage '+str(demBoost)+' Enthusiasm \n'
title2= str(demEnthus)+' Candidate Factor '+str(candidateFactor)+' Incumbancy ' + str(incumbent)+' Random '+str(spread)
plt.title(title1+title2,fontsize=200)
plt.axvline(218,linewidth=40,color='xkcd:bright green')
plt.axvline(194,linewidth=40,color='xkcd:red')

plt.savefig('C:\work\election2018\map5.png')


# In[7]:


# Democratic worst result
min(winList)


# In[8]:


# the percent chance of the Dems not taking the house
len([x for x in winList if x < 220 ])/sims




Saturday, November 14, 2015

Data Science - Live for the Learning Curve

I started doing data science in 1994. The tools I've used, in no particular order, are

 VSAM, JCL, SAS, SQL, PL/SQL, T-SQL, c-shell, Perl, R, Python, Java, Visual Basic, Tableau, Excel, flat files, Hadoop, HBase, Pig, Hive, DecisionSeries, AdminPortal

 That's a neat 21 tools in 21 years, and actually they are kind of obsolete already. There's a whole new data science paradigm starting of companies selling algorithms as APIs: send your data off in a web call, get a score back. Algorithms As A Service. Microsoft and Algorithmia come to mind.

 Whatever you're using now, wait a year: you'll have something new in your toolkit. If you're just getting through a data science course with R, Python, and Hadoop; well, that should keep you a couple of years.

 Live for the learning curve.

Sunday, August 16, 2009

Bozoing Measurements VII

A while ago I saw a consultant give a presentation. He had been given 20 campaigns to analyze. He spent a lot of time discussing the one campaign that was significant at the 5% level.

Sunday, July 12, 2009

Bozoing Campaign Measurements VI

And the hits keep coming.

This story involves a tracking database. The database was tracking long-running campaigns, where the process was that a customer 1) contacted the company via customer care 2) at that point, was randomized on a by-campaign basis. Once there was customer activity, that customer was tracked for three months.

Here's where it gets tricky. On the next customer contact the treatment group was given the pitch again if they still qualified whereas the control group was automatically not given the pitch. That means in the treatment group the next contact generates a campaign-relevant data point whereas in the control group it doesn't. Remember the three month-tracking? After three months any control group customers are dropped out of the database, whereas treatment group customers that are still in contact with the company are still tracked. These are long-running campaigns. So the control group was composed of customers that had at most a three-month window to take the offer whereas the treatment group had a potentially unlimited time to take the offer. What a clever way to make sure the results are excellent!

I was once reviewing analysis of campaigns from this system. I was originally asked to make sure the T-Test formula was right, and poked around in the data a little. I saw a weird thing: the campaign results were a linear function of the control group size. The smaller the control group the better the results. I commented that they really shouldn't publish results until they had figure out what the Weird Thing was. Looking back, I can see how the database anomaly aboce could account for the effect. As time goes on, customers are going to be dropped out of the control group. Also, the treatment group will be given longer and longer to take the offer. So as time goes on, the control group numbers will fall and treatment group takes will rise.

So all the positive results that were being ascribed to the marketing system could have been due to the reporting anomalies.

Saturday, June 27, 2009

Customers are Weird

Really, really weird.

Imagine a company with 2mm customers. Reasonable-sized, not huge.

How many people do you know well? Maybe 100 people? Think about the absolute weirdest person you know. That company has customers that are literally 100 times weirder than the weirdest person you know. In fact, they've got 200 of them.

It's a bad idea to think you know what customers are going to do without testing, measuring, and finding out.

Bozoing Campaign Measurements V

Here's a classic: toss all negative results.

Clearly, everything we do is positive, right?

Nope. Anything that can have an effect can have a negative effect.

(I've met a number of marketing people that really truly believe that people wait at home looking forward to their telemarketing calls. And that calling something 'viral' in a powerpoint is enough to actually create a viral marketing campaign).

There's another factor. Depressingly, a lot of marketing campaigns do absolutely nothing. Random noise takes over; half will be a little positive and half will a little negative. Toss the negative results and you're left with a bunch of positive results. Add them up and suddenly you've got significant positive results from random noise. This is bad.

I've seen an interesting variant on this technique from a very well-paid consultant. Said VWPC analyzed 20 different campaigns and reported extensively on the one campaign that had results that were significant at a 5% level.

Sunday, June 7, 2009

Bozoing Campaign Measurements - IV

Another installment in the "How to Bozo Simple Campaign Analysis". I've got a lot of them. It's amazing how inventive people get when it comes to messing up data.

Anyway, this is from a customer onboarding program. When the company got a new customer, they would give them a call in a month to see how things were going. There was a carefully held out control group. The reporting, needless to say, wasn't test and control. It was "total control" vs. "the test group that listened to the whole onboarding message". The goal was to enhance customer retention.

The program directors were convinced that the "recieve the call or not" decision was completely random; and given that it was completely random the reporting should be concentrated on only those that were effected by the program (that again -- it's amazing how often the idea comes up).

Clearly, the decision to respond to telemarketing is a non-random decision, and I have no idea what lonely neurons fired in the directors brains to make them think that. To start with, someone who is at home to take a call during business hours is going to be a very different population that people that go to work. More importantly, a person that thinks highly of a company is much more likely to listen to a call than someone who isn't that fond of a company.

Unsurprisingly, the original reporting showed a strong positive result. When I finally did the test/control analysis, the result showed that there was no real effect from the campaign.

Sunday, May 31, 2009

Statistics and DBAs

Statistics and DBA work really are two different disciplines, although from the outside we're both numbers people. I've learned the hard way that there's a lot that I don't know about how to set up a database. Likewise, I've had some database people push some very strange ideas about how to do analysis.

Take random samples. Unless I can actually see the code used to make random samples, I'd rather do random sampling myself. My favorite example of the problem was "we randomly gave you data from California".

Time sensitivity is another issue. I was making a customer attrition study for a cell phone company. We wanted to look at attrition over a year, so we needed customer data from the start of the year and we see how it effects attrition. What happened was that the database people, instead of following our instructions gave us customer data from the end of the year instead of start.

Why? "Don't you want the most current data possible?" It's the nature of reporting to get the most current data possible for the report, and understanding statistical analysis that will often require data from the past is a little alien to that way of thinking.

Bozoing Campaign Measurements - III

I've got another story from the customer cross-sell system I was talking about in Bozoing Measurements I

We're taking about doing basic reporting on the system. Remember, we're keeping out a control group. We were changing the control group process from keeping out individual control groups per each campaign (which caused a lot of problems actually -- more in a later post).

Now, the dead obvious comparison is treatment and control. There are a couple of nuances we can add on. We can compare

  • Total treatment vs. total control

  • The treatment and control that contact the company

  • The treatment and control that have an opportunity to be marketed to


All happy, all treatment vs. control.

But then the senior DBA in the project says "We shouldn't on report the control group that could be marketed to. That's a biased number".

Huh?

"That number is biased by the fact that we're taking out the customers that didn't contact us and that we couldn't market to."

His plan was to compare 1) Treatment group that contacted us and that we could market to (because the others clearly weren't effected by the program) to 2) The total control group. This would create a huge unfair effect favoring the treatment group, simply because the customers that are actively contacting the company are much more likely to purchase new products. That may have been the hidden agenda that the DBA had: create reporting that would have a large built in bias.

About that word bias: there's no such thing as a biased number. The number is what it is. Bias happens with unfair comparisons. We want the treatment and control factor to be the only factor in the comparison.

Tuesday, May 19, 2009

Bozoing Campaign Measurements -- II

Our next contestant comes from the telecom world.

What the group was doing was evaluating marketing campaigns over the course of several years. Does an attrition-prevention campaign have any effect after three years? This is an absolutely wonderful thing to do, of course, but not the way they went about it.

The campaigns were in a series of mailings that went out to customers that were about to go off contract, and the offer was a monetary reward to renew their contract for a year. Each campaign had a carefully selected control group.

The dead-obvious thing to do is to compare the treatment group vs. the control group, but that's not what got done. What happened was the analysis compared the whole control group to the customers in the treatment group that renewed their contract, because clearly "customers that didn't renew their contract weren't effected by the campaign".

Sound familiar?

Why doing analysis this way is a bad idea: before the mailing on contract renewal, customers are going to have a certain basic affinity towards the company. Some are going to love it, some are going to hate it, some are going to be on the fence. When the customers get the offer the ones that already hate the company will toss the offer, the ones that love the company will take free money for staying with a company they like, and the ones on the fence may or may not take the offer and have their future behavior change. So, to a good extent a retention program like this isn't changing behavior but instead is sorting the customers into buckets based on how they already feel about the company. Comparing "total control group" to "contract renewers" confounds two effects, one effect of the customers predisposition to the company and the second effect of having some customers renew their contracts for a reward. Moreover, this comparison doesn't actually answer the real question: does the program have a meaningful, measurable impact on churn? To answer the real question in the right way Keep Things Simple and Statistical and do a straight treatment vs. control.

Monday, May 18, 2009

How to Bozo Campaign Measurements

You know, at their heart statistical measurements are basically the easiest thing in the world to do, especially when it comes to direct marketing. Set up your test, randomly split the population, run the test, measure the results. It pretty much takes serious work to mess this up. It's amazing how many bright people leap at the chance to go the extra mile and find an inventive way to bozo a measurement.

The first exhibit is a database expert working for a customer contact project at a bank. A customer comes in, talks to the teller, and the system 1) randomly assigns the customer to the control group or not if this is the first time the customer has hit the system, otherwise it looks up the customer's status and then 2) makes a suggestion for a product cross-sell. The teller may or may not use the suggestion, depending on how appropriate the teller thinks the offer is for the customer and/or how busy the branch is and if there is time available to talk to the customer.

So now, we've got the simplest test/control situation possible. What the DBA decided was to toss out all the customers where no offer was made, on the theory that if no offer was made then the program had no effect. So, all the reporting was done on "total control group" vs. "treatment group that received the offer", creating a confounding effect. The teller decision to make the offer or not was highly non-random. The kind of person that comes in at rush hour (where the primary concern of the teller is handling customers and keeping wait times down) is going to be very different from the kind of person that comes during the slow time in the middle of the afternoon.

The project team understood this confounding, that in their reporting they were mixing up two different effects, and talked for over two years about how to overcome this confounding when all they had to do was be lazier and report on the random split.

Friday, February 13, 2009

The Data Daemon

Appropos of "Murphy's Laws of Data" , I find it useful to imagine that data is created by a little deamon and his job is to make me look like a durn fool.

Saturday, May 3, 2008

Learning from LTV at LTC: It's About Understanding

Ultimately, success is about understanding. Build teams that will take the time to understand the business and all parts of the project, where every member of the team understands all parts of the projects as a whole, share this understanding in full with anybody who wants to learn, and carry this detailed understanding forward in the enterprise.

Learning from LTV at LTC: Build Complete Teams

ypically projects are done by assembling cross-functional teams from different areas, each person with a narrow responsibility. This is a very efficient way of handling day-to-day business but an ineffective way of getting business-changing projects done. This is especially true is the project is going to be going on for a while.

The key to our success was having a complete team that could handle all phases of the project. There was no point in the project that we threw the project over the wall to another team, or caught something that another team was throwing at us. When we were working with other teams we established working relationships with them and brought those teams into the project. Every member on the LTV team could speak to all aspects of the project and have meaningful input into all aspects of the project.

Let me give an example of what can happen with fragmented, siloed teams. I was working on updating a project that had been launched several years before. There was one team that extracted the data from a datamart, another that took the data and loaded it into a staging area, and a third team that loaded the data from the staging area into the application. I asked the question “who can guarantee that the data in the application is right”? Thunderous silence. No one could guarantee that the final data was right, or even that their step was correct; all they could promise was that their scripts had run without obvious error.

If I had to give a name to this approach I'd call it the “A-Team” approach: complete functional teams that understand each other's areas.

Learning from LTV at LTC: Tell Everything

In a project like this the team gains a great deal of understanding about how the business works and there is always the temptation to keep that understanding within the team. The argument I have heard is that by keeping all the details hidden then the team will maintain control over the results of the project. What I've seen actually happen is that when a team tries to keep secrets others just don't believe them.

In the LTV project we made the decision to explain every detail to anybody who asked. The result was that people had a great deal of faith in what we produced. Even if people disagreed with the decisions that we made in the project, they understood and could respect the decisions.

Learning from LTV at LTC: Build Understanding

Projects that change an organization demand that the project group build a substantial understanding of that the business is, what it could be, and how the project can help the business get there. That understanding needs to stay withing the organization after the project is officially complete. There is a vast difference between the understanding that comes from seeing a presentation on a project and the understanding that comes from actually doing the work.

Projects that are important to the company need to be living, evolving things and that means that the detailed understanding of the project needs to stay accessible to the organization. With LTV, as soon as it came out people wanted additional work and we could do it because we knew the nuts and bolts.

Saturday, April 26, 2008

LTV at LTC: Learning from it: Design Rules

In software it's all about the implementation – actually writing the code. In business intelligence projects actually doing the implementation isn't that big a deal. There are lots of packages to make implementation easy compared to writing software from scratch. What that means is that business intelligence projects are all about the design, and the design team needs to be in control and actively involved in all stages of the project.

Thursday, April 24, 2008

LTV at LTC: The Large Activity Based Costing (ABC) Project

During and after the LTV project, there was yet a fourth value-based project at LTC. The Finance department brought in a large consulting company to design a database for activity-based costing to help LTC get a handle on their operational expenses. The goal was to build an ABC database where a manager could look at expenses, drill down into the specific line items, and then drill into the company and customer activity that was causing those expenses and so have a clear grasp of the actions needed to manage expenses.

The project started out by having the consultants come in and have roughly a year of large meetings on what should go into the system. This was done without considering implementation issues. At then end of the meetings a large and detailed specification was developed, which was then handed off to the LTC IT department. The LTC IT department estimated that implementation would cost several million dollars and the project was killed right then and there.

In many respects, the ABC project was the antithesis of the LTV project.

  1. Instead of identifying a group within the company to build the project, an outside consultant was brought in to run the project. This meant that the understanding that comes from doing a project like this left LTC with the consultants.
  2. There was a complete disconnect between the design and implementation teams. This meant that implementation issues were not considered during the design, and that the design could not be modified later to take implementation factors into consideration.
  3. Instead of a small group working to understand the business, ABC had large meetings to poll people on their issues. This meant that every possible issue was included in the project design. Because the design was simply thrown over a fence to implementation there wasn't any negotiation over project scope to achieve what was reasonable.