Friday, November 2, 2012

Predictions are Hard (Especially About the Future)

While famous malapropist Yogi Berra is most often cited for the quote, “Prediction is very difficult, especially about the future,” it appears that the source was actually Danish physicist and Nobel Prize laureate Niels Bohr. Bohr, whose pioneering work in quantum physics would naturally equip him with a keen sense of the limits of knowledge, also had a sense of humor. (He also said, “An expert is a man who has made all the mistakes that can be made in a very narrow field.”)

Bias and Accuracy

In my long study of cognitive biases on this blog and in my compilation Random Jottings 6: The Cognitive Biases Issue, I was struck again and again by how many of the biases had to do with perceptions of probability. From ambiguity aversion to the base rate fallacy to the twin problems of the gambler’s fallacy and the ludic fallacy, we have repeatedly shown ourselves to be incapable of judging probabilities with any degree of precision or understanding. When people rate their own decisions as "95% certain," research shows they're wrong approximately 40% of the time.

With the 2012 presidential election only four days away as I write this, the issue of prediction and forecasting is uppermost in the minds of every partisan and pundit. Who will win, and by how much? Checking the polls as I write, the RealClearPolitics average gives President Obama a 0.1% lead over Governor Romney (47.4% to 47.3%). Rasmussen has Romney up by 2 (49% to 47%), Gallup by 5 (51% to 46%), and NPR by 1 (48% to 47%). On the other hand, ABC/Wash Post and CBS/NY Times both have Obama leading by 1 (49% - 48% for ABC, 48% - 47% for CBS), and the National Journal has Obama up by 5 (50% - 45%). No matter what your politics, you can find polls to encourage you and polls to discourage you about the fate of your preferred candidate.

Some polls normally come with qualifications. Rasmussen traditionally leans Republican; PPP often skews Democratic. That doesn't means either poll is irrelevant or useless. Accuracy and bias are two different things. Bias is the degree to which a poll or sample leans in a certain direction. If a study comparing Rasmussen or PPP polls to the actual election results shows that Rasmussen's results tend to be 2% more toward the Republican candidate (or vice versa for PPP), both polls are quite useful — you just have to adjust for the historical bias. If on the other hand a poll overestimates the Democratic vote by 10% in one election and then overestimates the Republican vote by 10% in another election, there's no consistent bias, but the poll's accuracy is quite low. In other words, a biased poll can be a lot more valuable than an inaccurate one.

Selection Bias

Of course, political polls (or polls of any sort) are subject to all sorts of error. My cognitive biases entry on selection bias summarizes common concerns. For instance, there’s a growing argument that land-line telephone polls, once the gold standard of scientific opinion surveys, are becoming less reliable. Cell phone users are more common and skew toward a different demographic. There's also a sense that people are over-polled. More and more people are refusing to participate, meaning that the actual sample becomes to some extent self-selected: a random sample of people who like to take polls. People who don’t like to take polls are underrepresented in the results, and there’s no guarantee that class feels the same as the class answering. (I myself usually hang up on pollsters, and I've often thought it might help our political process if we agreed to lie to pollsters at every opportunity.)

Selection bias can happen in any scientific study requiring a statistical sample that is representative of some larger population: if the selection is flawed, and if other statistical analysis does not correct for the skew, the conclusions are not reliable.

There are several types of selection bias:

  • Sampling bias. Systemic error resulting from a non-random population sample. Examples include self-selection, pre-screening, and discounting test subjects that don’t finish.
  • Time interval bias. Error resulting from a flawed selection of the time interval. Examples include starting on an unusually low year and ending on an unusually high one, terminating a trial early when its results support your desired conclusion or favoring larger or shorter intervals in measuring change.
  • Exposure bias. Error resulting from amplifying trends. When one disease predisposes someone for a second disease, the treatment for the first disease can appear correlated with the appearance of the second disease. An effective but not perfect treatment given to people at high risk of getting a particular disease could potentially result in the appearance of the treatment causing the disease, since the high-risk population would naturally include a higher number of people who got the treatment and the disease.
  • Data bias. Rejection of “bad” data on arbitrary grounds, ignoring or discounting outliers, partitioning data with knowledge of the partitions, then analyzing them with tests designed for blindly chosen ones.
  • Studies bias. Earlier, we looked at publication bias, the tendency to publish studies with positive results and ignore ones with negative results. If you put together a meta-analysis without correcting for publication bias, you’ve got a studies bias. Or you can perform repeated experiments and report only the favorable results, classifying the others as calibration tests or preliminary studies.
  • Attrition bias. A selection bias resulting from people dropping out of a study over time. If you study the effectiveness of a weight loss program only by measuring outcomes for people who complete the whole program, it’ll often look very effective indeed — but it ignores the potentially vast number of people who tried and gave up.

Unskewing the Polls

In general, you can’t overcome a selection biases with statistical analysis of existing data alone. Informal workarounds examine correlations between background variables and a treatment indicator, but what’s missing is the correlation between unobserved determinants of the outcome and unobserved determinants of selection into the sample that create the bias. What you don’t see doesn’t have to be identical to what you do see. That doesn't stop people from trying, however.

With that in mind, the website unskewedpolls.com, developed by Dean Chambers, a Virginia Republican, attempts to correct what he sees as a systematic bias as to the proportion of Republicans and Democrats in the electorate. By adjusting poll results that in Chambers’ view are oversampling Democrats, he concludes (as of today) that Romney leads Obama nationally by 52% - 47%, a five point lead, and that Romney also leads in enough swing states that Chambers projects a Romney landslide in the electoral college of 359 to 179, with 270 needed for victory.

Chambers argues that other pollsters and analysts who show an edge for Obama are living in a “fantasy world.” In particular, he trains his disgust on Nate Silver, who writes the blog FiveThirtyEight on the New York Times website, describing him as “… a man of very small stature, a thin and effeminate man with a soft-sounding voice that sounds almost exactly like the ‘Mr. New Castrati’ voice used by Rush Limbaugh on his program. In fact, Silver could easily be the poster child for the New Castrati in both image and sound. Nate Silver, like most liberal and leftist celebrities and favorites, might be of average intelligence but is surely not the genius he's made out to be. His political analyses are average at best and his projections, at least this year, are extremely biased in favor of the Democrats.” (You may notice a little bit of ad hominem here. Clearly a short person with an effeminate voice can’t be trusted.)

A quick review of the types of selection bias above will identify several problems with the unskewed poll method. Indeed, it's hard to find anyone not wedded to the extreme right who's willing to endorse Chambers' methodology. The approach is bad statistics, and would be equally bad if done on behalf of the Democratic candidate.

Nate Silver and FiveThirtyEight

Other views of Nate Silver are a bit more positive. Silver first came to prominence as a baseball analyst, developing the PECOTA system for forecasting performance and career development of Major League Baseball players, then won some $400,000 using his statistical insights to play online poker. Starting in 2007, he turned his analytical approach to the upcoming 2008 election, and predicted the winner of 49 out of 50 states. This resulted in his being named one of the world’s 100 most influential people by Time magazine, and his blog was picked up by the New York Times. (He's also got a new book out, The Signal and the Noise: Why So Many Predictions Fail — But Some Don't. I recommend it.)

As of today, Nate Silver’s predictions on FiveThirtyEight differ dramatically from the UnSkewedPolls average. Silver predicts that Obama will take the national popular vote 50.5% to 48.4%, and the electoral college by 303 to 235. One big difference between Dean Chambers and Nate Silver is that Chambers is certain, and Silver is not. He currently gives Obama an 80.9% chance of winning, which means that Silver gives Romney a 19.1% chance of victory using the same data.

This 80% - 20% split is known to statisticians as a confidence interval, a measure of the reliability of an estimate. In other words, Silver knows that the future is best described as a range of probabilities. Neither he, nor Chambers, nor you, nor I “know” the outcome of the election that will take place next Tuesday, and we will not “know” until the votes have been counted and certified (and any legal challenges resolved).

Predictions vs. Knowledge

In other words, when we predict, we do not know.

Keeping the distinction straight is vital for anyone whose job includes the need to forecast what will happen. Lawyers don’t “know” the outcome of a case until the jury or judge renders a verdict and the appeals have all been resolved. Risk managers don’t “know” whether a given risk will occur until we’re past the point at which it could possibly happen. Actuaries don’t “know” how many car accidents will take place next year until next year is over and the accidents have been counted. But lawyers, risk managers, actuaries — and pollsters — all predict nonetheless.

A statistical prediction, by its very nature, contains uncertainty and should therefore be expressed in terms of the degree of confidence that the forecaster has determined. “The sun’ll come out tomorrow,” sings Annie in the eponymous musical, and she’s almost certainly right. But that’s a prediction, not a fact. While the chance of the Sun going nova are vanishingly small, they aren’t exactly zero.

Confidence Level and Margin of Error

Poll results usually report both a confidence level and a range of error, such as “95% confidence with an error of ±3%.” The error rate is the uncertainty of the measurement itself. If we flip a coin 100 times, the theoretical probability is 50 heads and 50 tails, but if it came out 53 heads and 47 tails (or vice versa), no one would be surprised. That’s equivalent to an error of ±3%. In other words, a small wobble in the final number should come as a shock to no one.

The confidence level, on the other hand, is the degree of confidence you have that your final number will stay within the error range. The probability that an honest coin flipped 100 times would produce 70 heads and 30 tails is low, but it’s within the realm of possibility. In other words, the “95% confidence” measurement tells us that 95% of the time, the actual result should be within the margin of error — but that 5% of the time, it will fall outside the range. (There’s a bit of math that goes into measuring this, but it's outside the scope of this piece.)

Winning at Monte Carlo

Nate Silver’s 80% confidence number comes from using a modeling technique known as a Monte Carlo simulation, which is also used in project management as a modern and superior alternative to the old PERT calculation, a weighted average of optimistic, pessimistic, and most likely outcomes. In a Monte Carlo simulation, a computer model runs a problem over and over again in thousands of iterations, choosing random numbers from within the specified ranges, and then calculates the result. If the polls are right 95% of the time within a ±3% margin of error, the program chooses a random number within the error range 95% of the time, and 5% of the time chooses a number outside the range, representing the probability that the polls could be all wet. In running five or ten thousand simulations, the results gave the victory to Obama 80.9% of the time, and to Romney 19.1% of the time.

Tomorrow, the answer may be different. Silver will enter new data, and the computer will run five or ten thousand more simulations. Each day, the probability of winning or losing will change slightly, until the final results are in and the answer is no longer a matter of probability but a matter of fact.

The Thrill of Victory and the Agony of Defeat

Astute readers may notice the parallels here to Schrödinger's Cat, which is mathematically both alive and dead until the box is opened. Personally, I put a lot of credence into Silver’s analysis; his approach is in line with my understanding of statistics. That means I think Obama is very likely to win next Tuesday — but only within a range of probability.

I will also note that Nate Silver seems to feel the same way. He's just been chided by the public editor of the New York Times for making a $2,000 bet with "Morning Joe" Scarborough that Obama will win. Given his estimate of an 80% - 20% chance of an Obama victory, that sounds like a pretty good bet to me.

But we won't know until Tuesday night at the earliest. So be sure to vote.

Saturday, October 20, 2012

Saturday Night's Alright (for Firing) — Watergate, Part 9

Watergate Special Prosecutor
Archibald Cox
For previous installments of my irregular series tracing the history of the Watergate scandal, click here. This week, the Saturday Night Massacre, October 20, 1973.

The Watergate burglary itself took place on June 20, 1972, but following Richard Nixon's overwhelming re-election in November of that year, it looked as if the worst of the scandal had been contained. As long as the burglary could be put down to overzealous underlings at the Committee to Re-Elect the President (CRP, but often abbreviated CREEP) and kept away from the White House itself, all was in order.

There were loose ends. One of the burglars had checks from E. Howard Hunt, a member of the White House "plumbers" who was connected to Special Counsel to the President Charles Colson, known as Nixon's hatchet man. As part of the cover up, White House Counsel John Dean went to acting FBI director L. Patrick Gray to keep the situation under control. As Dean later wrote, "[We] could count on Pat Gray to keep the Hunt material from becoming public, and he did not disappoint us."

Gray went so far as to burn what were billed as "national security documents [that] should never see the light of day" from Hunt's personal safe at the request of Dean and Assistant to the President for Domestic Affairs John Ehrlichman. These documents weren't officially about Watergate, Gray later said. "The first set of papers in there were false top-secret cables indicating that the Kennedy administration had much to do with the assassination of the Vietnamese president (Diem). The second set of papers in there were letters purportedly written by Senator Kennedy involving some of his peccadilloes, if you will."

Unfortunately, Gray wasn't the only person who knew about the Hunt material. His deputy, FBI Associate Director W. Mark Felt, who actually ran the FBI's day-to-day operations, was also "Deep Throat," the confidential informant providing Washington Post reporters Bob Woodward and Carl Bernstein with information. As the material began to leak, Gray became shaky.

In February 1973, Nixon nominated Gray to be permanent director of the FBI, handing the Senate its first opportunity to interrogate a high-ranking Administration official about Watergate. Gray went into full self-defense mode. He volunteered that he'd provided investigation files to John Dean, saying FBI lawyers had told him it was legal, confirmed the dirty tricks activities of CREEP — and worst of all, testified that Dean himself had "probably lied" to the FBI. Enraged by the betrayal, Ehrlichman told Dean that Gray should "twist slowly, slowly in the wind." (Ehrlichman was evidently a fan of Huxley's Brave New World.) Gray withdrew his nomination, and after he learned that Dean had rolled over, Gray resigned from the FBI altogether. Although he was later indicted, he was never convicted.

In March 1973, Watergate burglar and CREEP security specialist James McCord wrote Watergate Judge John Sirica that his testimony was perjured under pressure. One month after that, seeing the handwriting on the wall, John Dean rolled over and began cooperating with Federal prosecutors. Desperate to distance himself from the scandal, Nixon responded by firing Ehrlichman, White House Chief of Staff H. R. Haldeman, and Attorney General Richard Kleindienst. (Kleindienst had taken over from John Mitchell when Mitchell was tasked with leading the re-election effort. His involvement with the scandal was peripheral, and he ended up with a misdemeanor conviction for perjury and paid a $100 fine.)

With the Justice Department compromised, Nixon had little choice but to allow the appointment of a nominally independent special prosecutor, Archibald Cox. After the revelation of the White House tapes and Nixon's refusal to release them, Cox pursued a subpoena to get the tapes for his investigation. When Cox refused a Nixon compromise that would give him transcripts but no access to the actual recordings, Nixon had had enough.

On Saturday evening, October 20, 1973, Nixon called Attorney General Elliot Richardson, Kleindienst's successor, and ordered him to fire Cox. Richardson, citing his promise to the Congressional oversight committee not to interfere with the Special Prosecutor, refused.  When Nixon continued to press him, he resigned. Nixon then called the Deputy Attorney General, William Ruckelshaus, who had made the same pledge, and ordered him to fire Cox. Ruckelshaus also resigned.

The third in command of the Justice Department was Solicitor General Robert Bork (later a notorious failed Supreme Court nominee), who had not been part of the process and who had therefore not made the same pledge. Although Bork claimed to believe that Nixon had the right to fire Cox, he says he also considered resigning so he wouldn't be "perceived as a man who did the President's bidding to save my job." Elliot Richardson says he persuaded Bork not to resign, on the grounds that the Justice Department needed some continuity of leadership.

Nixon had Bork brought to the White House by limousine, swore him in as Acting Attorney General, and had Bork write the letter on the spot firing Cox.

This incident became known as the "Saturday Night Massacre," and it was a major tipping point in the scandal. Congress was infuriated, the public outraged. After the Massacre, a plurality of Americans for the first time supported impeachment: 44% for, 43% against, 13% undecided. Several resolutions of impeachment were introduced in the House. Nixon was forced to allow Bork to appoint a new special prosecutor, Leon Jaworski. There was some concern Jaworski, as the President's approved choice, would limit the investigation to the burglary alone, but as it turned out, Jaworski also looked at the broader implications of the growing scandal.

In November 1973, a Federal district judge ruled that Cox's firing was illegal under the regulation establishing the special prosecutors office, which required a finding of "extraordinary impropriety." However, the situation had moved far too quickly to allow Cox to resume his position. The battle of the tapes would continue well into the following year.


Tuesday, October 2, 2012

Fifty Thousand!

 

I was pleased to discover yesterday that my Sidewise Thinking blog has now hit the 50,000 pageview mark. Last month, there were over 4,300 views, or well over 150 per day.

My first post, "What's SideWise Thinking?", appeared on April 11, 2009. It was an excerpt from the book I was currently working on, Creative Project Management (with Ted Leemann). I've generally put a new post up every Tuesday (with a big gap between July and November 2010), with topics ranging from project and risk management to my two big series on cognitive biases and decision-making disorders.

The most popular piece so far has been "You're Not Being Reasonable," on the rules of reasonable arguing. First published on March 2, 2010, it's gotten over 3,400 page views, helped primarily by a plug from the blog "LessWrong" and a StumbleUpon link.

I don't quite understand why the second most popular post is the 23rd part of my Red Herrings series, "Hume's Guillotine." First published January 24, 2012, it's gotten over 2,200 hits, but I can't find any specific factor driving traffic to that article and that one alone. Next comes "Triage for Project Managers (Part Two)" (February 8, 2011, over 1,700 hits), and "Eyewitness to Murder" (April 13, 2010, with over 1,300). Red herrings strike again with "A Cute Angle (Part 19)" (December 27, 2011, over 1,000 hits).

By comparison, my new blog, Dobson's Improbable History, which has only a little more than a month under its belt, is already exceeding 100 hits per day, with over 3,200 pageviews last month — a much better start.

This is the 149th post I've made to the blog. I made 29 entries in 2009, 31 in 2010, 52 in 2011, and 43 so far this year.

Thanks very much for reading, and I hope you continue to enjoy it.




Tuesday, September 25, 2012

Goldfinger Takes Fort Knox! (Propositional Fallacies, Part 2)

Bond villain Auric Goldfinger
In propositional calculus, we can describe certain arguments in mathematical terms. Some arguments are true if the component statements are true. The statement “It is raining here now, and it is raining where you are now as well” can be written as P⋀Q. It is true if both its component statements are true. On the other hand, “It is raining here now OR it is raining where you are now” (written as P⋁Q) is true as long as at least one of the statements is true.

Propositional fallacies involve fallacies of mathematical reasoning. They are fallacious regardless of the truth value of the component statements. Last time, we discussed affirming a disjunct, the fallacy of turning an inclusive OR into an exclusive one. The two remaining propositional fallacies are known as affirming the consequent and denying the antecedent.

Affirming the Consequent

If Auric Goldfinger owned Fort Knox, then he would be rich. Auric Goldfinger is rich. Therefore, Auric Goldfinger owns Fort Knox. Even if the first two statements are true, the conclusion is invalid because there are other ways to be rich besides owning Fort Knox.

Here's how to cast the argument in propositional calculus:

P→Q 
∴ P 

(If P, then Q. Q is true. Therefore, P.)

This is different from the argument "if and only if." If Auric Goldfinger is rich if and only if he owns Fort Knox, then the statement "Auric Goldfinger is rich" makes "Auric Goldfinger owns Fort Knox" necessarily true. But that's the case only if the first statement is true — which it isn't. In propositional calculus, we'd write that:

P⟷Q
Q
∴ P

Affirming the consequent is sometimes called converse error.

Denying the Antecedent

The opposite fallacy, denying the antecedent, is also known as inverse error.

If Auric Goldfinger owned Fort Knox, then he would be rich. Auric Goldfinger does not own Fort Knox. Therefore, Auric Goldfinger is not rich. This is wrong for the same reason as the previous argument was wrong: there are other ways to be rich.

In propositional calculus, this takes the form:

P→Q 
 ¬P
∴ ¬Q

If P, then Q. P is false (not-P). Therefore, Q is false (not-Q). As in the previous case, the rules for if and only if are different from if alone.



Monday, September 17, 2012

The Seven Deadly Sins — and Where To Find Them

Researchers at Kansas State University decided to create a series of county-by-county maps of the United States showing the relative distribution of the Seven Deadly Sins (Envy, Greed, Wrath, Sloth, Gluttony, Lust, and Pride). For each sin, they identified a measurable criterion that could serve as a stand-in, and mapped the results showing the deviation from the norm expressed in terms of the standard deviation (σ). Measures from -1.65σ to + 1.65σ are normal; lower levels shade toward the blue and higher levels toward the red.

It's very easy to critique the criteria used for each sin, or to suggest alternative metrics, but I thought it was quite interesting nonetheless. You can learn more about the project and the researchers here, starting on page 8 of the PDF.

Envy

Metric: Total thefts (robbery, burglary, larceny, grand theft auto) per capita.



Maps of the Seven Deadly Sins



Gluttony

Metric: Number of fast food restaurants per capita.







Greed

Metric: Average income compared with the number of people living below the poverty line.




Lust

Metric: Number of STD cases reported per capita.





Sloth

Metric: Expenditures on art, entertainment, and recreation compared with employment.





Wrath

Metric: Number of violent crimes (murder, assault, rape) per capita.




Pride

Metric: Aggregate of the other six offenses — because pride, as they say, is the root of all sin.






Tuesday, September 11, 2012

Propositional Fallacies, Part 1


There’s a branch of math known as propositional calculus that treats arguments like mathematical propositions. Using propositional calculus, you can demonstrate the truth or falsity of certain arguments.

Take the statement “It is raining here now.” Depending on when you make the statement, it can be either true or false. In propositional calculus, you’d represent the statement as “P,” and the opposite, “It is not raining here now” as “¬P.” If P is true, then ¬P has to be false; if ¬P is true, then P has to be false.

You can link together statements with connectors. Common connectors are AND, OR NOT, ONLY IF, and IF AND ONLY IF. If we say “It is raining here now, and it is raining where you are now as well,” we can label the second statement as Q. Represent AND with the symbol ⋀, and we can write “It is raining here now, and it is raining where you are now as well” as P⋀Q.

Of course, maybe it is raining here or it isn’t; maybe it’s raining at your house and maybe it isn’t. Because the individual statements can be true or false, we can prepare a truth table.

P                    Q                    P⋀Q
True         True            True
True         False              False
False             True            False
False             False           False

With and as a connector, the proposition P⋀Q is only true if both statements are true.

The connector OR (represented as “⋁”), on the other hand, makes the proposition true as long as at least one of the statements are true. “It is raining here now OR it is raining where you are now” results in the following truth table.

P                    Q                    P⋁Q
True         True           True
True         False          True
False          True              True
False           False             False

Notice that OR is used here inclusively rather than exclusively. That is, P doesn’t exclude Q from being true. If it’s raining at my house, that doesn’t mean it’s not raining at yours.

Given the idea of propositional logic, it's easy to conclude that there are fallacies to go with it. The first of these is known as affirming a disjunct.

Affirming a Disjunct

Also known as the fallacy of the alternative disjunct, or the false exclusionary disjunct, this particular fallacy occurs when you change an inclusive OR into an exclusive one. “It is raining here now or it is raining where you are now” gets interpreted as “If it is raining here now, then it isn’t raining where you are now.”

In our symbolic structure, that gets represented as the following argument (with “therefore” represented by ∴).

P⋁Q
P
∴¬Q

That’s a fallacy because it could be raining both places. One doesn’t preclude the other.

While OR in logic always means an inclusive “or,” that doesn’t mean you don’t sometimes want to be more concrete. The logical operator XOR is an exclusive or. When you use it, you’re saying “one or the other, but not both.” The symbol for that is ⊻.

More next week.

Tuesday, September 4, 2012

Who Was That Masked Man? (Formal Fallacies Part 3)


Formal fallacies are arguments that are always wrong, regardless whether the argument's premises (statements claimed as fact) are true or false. For example, in the appeal to probability, someone makes a claim that because something could happen, therefore it will happen. That’s false even if it's true that the something in question could indeed happen.

Masked Man Fallacy

I know who Bruce Wayne is.

I do not know who Batman is.

Therefore, Bruce Wayne is not Batman.

In the masked man fallacy, a substitution of identical designators in a true statement can lead to a false one. The statement "I do not know who Batman is" gets treated as if it excludes Bruce Wayne simply because I do know who he is. Of course, as long as I don’t know that Bruce is actually Batman, both statements can be absolutely true, and yet the conclusion does not follow logically.

The general form of the argument is:
X is known.
Y is unknown.
Therefore, X is not Y.
A similar argument, however, is valid.

Clark Kent is Superman (X is Z).

Batman is not Superman (Y is not Z).

Therefore, Clark Kent is not Batman (therefore, X is not Y).

That’s because being something is different from knowing something. Lack of proof of one proposition doesn’t serve as proof of the counter proposition.