TruthIsAll
03-26-2009, 10:50 AM
In 2006, 120 Pre-election House Generic Polls asked the question:
Which congressional candidate are you going to vote for, the Democrat or Republican?
I created a trend line model based on a simple linear ("best-fit") regression of the 120 poll shares.
The model projected the Democrats would win 56.43% of the House Generic vote.
The Linear Regression Trend Model produced the following "best fit" trend lines.
Dem = 46.98 + .0419x
Rep = 38.06 + .0047x
Substituting x = 120 and allocating 60% of undecideds (UVA) to the Democrats:
..... Trend + UVA = Projection
Dem = 52.01 + 4.42 = 56.43%
Rep = 38.62 + 2.95 = 41.57%
1http://img.photobucket.com/albums/v474/autorank/Election2006_16921_image001.png
Exit pollsters Edison-Mitofsky asked 13,251 responders:
Who did you just vote for, the Democrat or Republican?
http://www.ropercenter.uconn.edu/elections/common/exitpolls.html
6351 respondents were asked who they voted for in 2004
The unadjusted exit poll of 13251 respondents indicated that the Democrats won the generic vote by:
Dem Rep
7470 5476
56.37% 41.33%
The pre-election Generic 120 poll trend matched the exit poll to within .07%!
Scenario 1:
Assume the unadjusted exit poll mix of returning Kerry and Bush voters:
Adjust vote shares to match the unadjusted exit poll.
Result:
The Final NEP was forced to match the 52.2-45.9% recorded margin by
1) adjusting the returning Kerry/Bush voter mix to an implausible 43/49%
2) lowering the Democratic shares of returning Kerry and Bush voters by approximately 2-3%.
In order to match the 56.37-41.33% vote split using 7pm NEP vote shares, the returning voter mix had to be revised:
1) the impossible 4.0% Other mix was lowered to 1.0% to match the 2004 recorded share.
2) the Kerry/Bush returning voter mix was changed to 48.6/46.4%.
Bush vote shares had to be lowered to match the unadjusted exit poll.
Adjusted 7pm NEP shares Final Final NEP shares
MIX Dem Rep Other MIX Dem Rep Other
Kerry 48.6% 93% 6% 1% 43% 92% 7% 1%
Bush 46.4% 17% 81% 2% 49% 15% 83% 2%
Other 1.0% 65% 18% 15% 4% 66% 23% 11%
DNV 4.0% 66% 16% 19% 4% 66% 32% 2%
TOTAL 100.00% 56.38% 41.32% 2.32% 100% 52.19% 45.88% 1.93%
Margin 15.06% 6.31%
This is further evidence that Kerry must have won in 2004.
______________________________________________________________________
The Final NEP radically changed the 7pm MIX and shares in order match the recorded vote.
Note that even the 7pm returning voter mix was highly dubious since it assumes that
the 2004 recorded vote was the True Vote (we know it wasn't)
7pm NEP Final NEP
MIX Dem Rep Other MIX Dem Rep Other
Kerry 45% 93% 6% 1% 43% 92% 7% 1%
Bush 47% 17% 82% 1% 49% 15% 83% 2%
Other 4% 66% 24% 10% 4% 66% 23% 11%
DNV 4% 67% 30% 3% 4% 66% 32% 2%
TOTAL 100% 55.16% 43.40% 1.44% 100% 52.19% 45.88% 1.93%
Margin 11.76% 6.31%
___________________________________________________________________________________
Scenario 2:
Assume that
1) The returning voter mix was proportional to the 2004 exit poll (Kerry 52-47%)
2) Returning Other voters were 1% of the 2006 total
3) Preliminary 7pm NEP vote shares
The resulting Democratic vote share is 57.87%.
In order to match the recorded vote, the Final NEP had to be adjusted by:
1) reducing the 7pm NEP Democratic vote shares
2) reversing the returning mix from +6% Kerry to +6% Bush
a) increasing the Other mix from 1% to 4% by deducting 3% from Kerry.
b) increasing the Bush mix by and decreasing the Kerry mix.
Unadjusted NEP Final NEP
MIX Dem Rep Other MIX Dem Rep Other
Kerry 50.5% 93% 6% 1% 43% 92% 7% 1%
Bush 44.5% 17% 82% 1% 49% 15% 83% 2%
Other 1% 66% 24% 10% 4% 66% 23% 11%
DNV 4% 67% 30% 3% 4% 66% 32% 2%
TOTAL 100% 57.87% 40.96% 1.17% 100% 52.19% 45.88% 1.93%
_____________________________________________________________________________
Now for a little history. In 2006, I posted the results of the 120 Generic Poll Trend model on Scoop, DU and PI (Dem 56.4 - Rep 41.6%).
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=364x2600324
But the Final NEP (Dem 52-Rep 46%) was way off (as usual, it was forced to match the recorded vote):
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=203x459299
So I posted an analysis which calculated a near zero probability that the recorded vote would deviate by over 4% from the 120 poll trendline.
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=364x2775205
Immediately, the usual DU naysayers jumped in. Surprisingly, Skinner joined them. He offered a lengthy polemic in which he lambasted my use of pre-election Generic polls to project the House vote. As far as I know, Skinner was not experienced in statistical polling analysis (he never commented in the pre-election and exit poll debates which permeated DU in 2004-2005. Apparently, he became an expert overnight. And now he was going to teach DUers all about Generic polls - and bash my "rookie" analysis in the process.
He claimed that pre-election Generic Polls should not be used for projecting the vote. He used caps for emphasis (I used to use caps a lot, too): It is "WRONG, WRONG, WRONG". He admitted that he didn't check my mathematics, but conjectured that I may have also made errors in the calculations.
He called my analysis an "embarrassment". He told DUers to THINK.
He ended with this advice: "You can do the best, most accurate, most awesome mathematics in the history of the world, but if you start with completely false assumptions, your "analysis" is going to be worthless".
Apparently, Skinner was unaware of the fact that for 40 years I have developed mathematically based scientific and financial models for many of the world's largest defense, consumer and financial corporations. Or that I developed and marketed corporate financial software to major banks and consumer product manufacturers. I assume Skinner is still in his thirties, so I was doing this before he was born.
_____________________________________________________________________
Despite the fact that the pre-election generic polls (adjusted for the 60% Democratic share of undecided voters) matched the 7pm national exit poll, Skinner, the new self-proclaimed polling "expert", wrote the following:
To those of you who keep demanding to see the problem, here it is:
http://www.democraticunderground.com/discuss/duboard.php?az=show_mesg&forum=364&topic_id=2775205&mesg_id=2781808
The problem is not in the mathematics (although I have not checked the math, so it's possible that there are errors there, too). The problem is in the assumptions he used before he even started.
TIA assumes that the "generic poll" should match the recorded votes. This is, quite simply, WRONG WRONG WRONG WRONG WRONG.
Does everyone here know what the "generic poll" is? The generic poll (usually referred to as the "generic congressional ballot") asks respondents which political party they support in the upcoming congressional election -- but it does not provide any names of any candidates. The generic congressional ballot is not -- and was never intended to be -- an accurate prediction of how people will vote. The point of the generic congressional ballot is to get a general sense of the mood of the voters.
Think, people. THINK.
Most American voters are not political junkies or activists. Even during an election season they could not tell you for certain the name of their member of Congress. When they get a call asking them their preference between a generic Democratic candidate and a generic Republican candidate, they will simply respond based on party. A large proportion of voters would be unable to think of the name of the candidate that represents each party in their upcoming congressional election.
But what happens when those same voters enter the voting booth on election day? They are presented with a list of NAMES OF REAL PEOPLE, their candidates, listed by office and political party. And they remember what they like and dislike about all of these people. They might have expressed support for a particular party, and they might hold that generic opinion of that party, but they like their own member of Congress and don't particularly care what party he or she is in. Or they might remember the ads about how a particular candidate was a crook or sexual deviant, and now that they finally face the realization he or she is running in their congressional district, they cannot possibly vote for him or her regardless of what party he represents.
Bottom line:
You can do the best, most accurate, most awesome mathematics in the history of the world, but if you start with completely false assumptions, your "analysis" is going to be worthless.
If you want to see how the outcome compares to the pre-election polls, you need to look at the pre-election polls that list candidates by name from each and every congressional district. Someone mentione up-thread that this is precisely what folks like Charlie Cook and Stu Rothenberg did before the election, and their predictions were quite accurate.
________________________________________________________________
I responded to that drivel with this post:
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=203x460141
I provided examples to show that the Generic polls ALL asked the proper question:
FOX News/Opinion Dynamics Poll. Nov. 4-5, 2006.
N=900 likely voters nationwide. MoE ± 3.
"Thinking ahead to this November's elections, if the congressional election were held today, would you vote for the Democratic candidate in your district or the Republican candidate in your district?" If unsure: "Well, if you had to vote, which way would you lean?"
I reviewed the Generic poll projection model:
The projected Democratic vote share was based on the trend line of ALL 120 Generic polls taken from Sept. 2005 up to Election Day. The Democrats won ALL 120 polls by an average 13.24% margin, 51.84D-38.60R, with 2% going to 3rd party candidates. I allocated 60% (UVA) of the 7.56% undecided voters to the Democrats. The final projection was 56.43D-41.62R, a 14.81% margin.
I explained the rationale for using a 1.5% MoE to compute the probability of the discrepancy:
One DUer made the following statement regarding the analysis: "Nowhere does it show how that polling sample would be 25% less variable than a standard poll, which almost always round DOWN (due to sample size) to 2%. That would inflate the distribution of the final analysis by a factor of 4/3 which moves the tails farther from the mean. Since the distributions are not linear, at that end of the curve, it could move the probabilities out by a factor of more than 100".
Here's why the 1.5% MoE was justified: I used the FINAL COMBINED 10 polls to calculate the MoE. There were 1000 sampled in each poll. Ten (10) INDEPENDENT Final Generic Polls are essentially equivalent to ONE poll of 10,000 sample size. The MoE for a 10,000 sample is near 1.0%. So the 1.5% MoE assumption was a conservative one.
The formula used to calculate MoE is:
MoE = 1.96*standard error = 1.96*SQRT((1-p)*p)/n)
For p=.56 and n= 10,000 sample-size, MoE = 0.97%.
I reviewed Generic polls and basic statistics:
Generic polls are designed to sample representative congressional districts. They all ask the same question. Calculating an average trend line or arithmetic mean gives us greater confidence that the sample mean is close to the true population mean. Is there anyone who will question the Law of Large Numbers and the Central Limit Theorem?
_________________________________________________________________________________
Skinner responded to the OP with a non-response. He obviously realized that he was just called out:
I'll save myself the effort and link to my previous post.
http://www.democraticunderground.com/discuss/duboard.php?az=show_mesg&forum=364&topic_id=2775205&mesg_id=2781808
Because my response is still correct, and TIA's central assumption is still wrong. He can try to dazzle people with a wall of words and numbers, but it will not matter. He cannot turn falsehood into truth simply because he wishes it to be so.
_________________________________________________________________________________
anaxarchos replied to Skinner:
You might want to exert that effort...
He proceeded to provide evidence that reputable pollsters and political scientists use Generic polls in their projections:
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=203x460141#460379
The central thesis in your link is summed up by this:
"If you want to see how the outcome compares to the pre-election polls, you need to look at the pre-election polls that list candidates by name from each and every congressional district. Someone mentioned up-thread that this is precisely what folks like Charlie Cook and Stu Rothenberg did before the election, and their predictions were quite accurate."
Even a casual search on Charlie Cook and Stu Rothenberg and "generic polls" yields dozens of both pre-election and post-election comments by both using generic polls in a similar manner to TIA. Certainly, they checked the generics against specific contests and other polling but, in this election, the generics seem to have been the most important instrument used by both and this, both in their detailed projections and in their post-mortems.
There have been problems with generics in the past, but, they were clearly useful in 2006 and neither Cook nor Rothenberg is a very good source for your blanket dismissal. Given that, you may want to explain why a compilation of over one hundred generics constitutes "completely false assumptions" or a "worthless" analysis.
http://www.cookpolitical.com /
November 6, 2006
All Monday there was considerable talk that the national picture had suddenly changed and that there was a significant tightening in the election. This was based in part on two national polls that showed the generic congressional ballot test having tightened to four (Pew) and six (ABC/Wash Post) points.
(snip)
Furthermore, there is no evidence of a trend in the generic ballot test. In chronological order of interviewing (using the midpoint of field dates), the margins were: 15 points (Time 11/1-3), 6 points (ABC/Wash Post), 4 points (Pew), 7 points (Gallup), 16 points (Newsweek), 20 points (CNN) and 13 points (Fox).
In individual races, some Republican pollsters see some movement, voters "coming home," in their direction, and/or some increase in intensity among GOP voters. All seem to think that it was too little, too late to significantly change the outcome. However, it might be enough to save a few candidates. None think it is a major change in the dynamics of races, and most remain somewhere between fairly and extremely pessimistic about tomorrow's outcome.
http://blog.washingtonpost.com/thefix/2006/08/parsing_t...
The answer, according to Charlie Cook and Stu Rothenberg, is a guarded yes.
"If you take an average of the last three or four polls, because any one can be an outlier in either direction, you can determine which way the wind is blowing, and whether the wind speed is small, medium, large or extra-large," said Cook. "The last three generics that I have seen have been in the 18 or 19 point range, which is on the high side of extra large. That suggests the probability of large Democratic gains."
"The generic surely reflects voters dissatisfaction with the President and his party and their inclination to support Democrats in the fall," agreed Rothenberg. "The size of the Democrats' generic advantage also can't be ignored. It too suggests the likelihood of a partisan wave, even though it does not guarantee the fate of any individual Republican incumbent."
http://rothenbergpoliticalreport.blogspot.com/2006/07/w...
Second, the ballot test in the July 10-13 Cooper & Secrest poll strongly mirrors the generic ballot in the district. Donnelly leads Chocola by 10 points in the ballot test (48 percent to 38 percent) while a generic Democratic candidate leads a generic Republican by 10 points as well (46 percent to 36 percent). The Democrats’ generic ballot advantage grew from 1 point in November 2005 to 10 points earlier this month, which also helps explain Donnelly’s improved standing in the July survey.
Personally, I'm not sure about the probability or extent of fraud in 2006... or in the practicality of mass fraud in mid-term congressional elections in general. Nevertheless, dismissing that possibility is not as easy as we might want it to be.
_________________________________________________________________________________
Skinner weakly replied to anaxarchos:
107. They're using the generic ballot in a similar manner to TIA? Oh really?
Find me the place where they use generics to argue that millions of Democratic votes were lost. Or find me a place where they claim the generic ballot provides an accurate prediction of the final vote count.
Do they use generics to inform their predictions? Of course they do. Generic ballots are a useful tool. I have never argued that they are not, and I have never argued that they are completely worthless. What I have said is that they are not supposed to accurately predict the final vote. All of the quotes you have provided in your post are fully consistent with my argument.
And, for the record, I am not "dissmissing the possibility of fraud." I consider it offensive and misleading to suggest that I have some sort of hidden nefarious agenda. I fully support any and all efforts to secure our elections and to prevent and uncover fraud. But that doesn't mean that I am required to shut off my brain and pretend not to notice when people on my own side make glaring errors. TIA may not care, but in the real world credibility matters. You can't keep making very basic rookie mistakes without people eventually dismissing the entire effort as a fabrication.
_________________________________________________________________________________
Sancho, a very well-respected FL Election Supervisor, commented:
167. A good point by TIA....from the scholars!
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&
address=203x460141#461262
As I've been learning about the world of pollsters (thanks to links from Febble and my own searching), I agree with much that I found in these articles from a single volumn of POQ. The exit pollsters have sacrificed accuracy for a single poll that is fast, but has too much error. TIA is appropriately using multiple sources that serve as checks and balances against each other, or multiple sources of evidence.
It seems to me, that many of the criticisms of the VNS exit polls are intuitively as serious as those committed by TIA, even if TIA's "math" and "assumptions" are not peer reviewed! Mitofsky (if you read between the lines) pretty much describes the poll failures (at that time) in "technical terms", and the connection of pre to exit is a suggested solution, as is the use of multiple polls, as is appropriate sample sizes, etc. In effect, VNS called elections with logic mistakes at least as "bad" as TIA is accused of on this blog, but the VNS had the responsibility and paycheck to get it right!
The pollsters also refuse to consider fraud (as alleged here on occasion), but admit they cut corners. The original poll designers also describe TV station analysis performing mathematical projections (comparison with previous year's polls) without the power and data to do the job, essentially not meeting the assumptions. Because non-response error is common and often cited doesn't mean that the more exotic, but equally bad mistakes by the VNS don't deserve attention; and considering fraud as a source of variation is clearly missing from the "scholars" as often as TIA suggests it as the problem.
_________________________________________________________________________________
Sancho replied to Febble:
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=203x460141#460182
TIA's formula "mistake" affects power as much as anything. Regardless, even the "odds" that 100 to 1, much less 100,000,000 to 1 that an election is hacked is pretty serious. I think that political correctness among pollsters is to avoid making the accusation, and I now accept that...but what can be done by pollsters to help with post-hoc analyses that would demand a different system or revote? Florida judges will USE the POLLSTER'S conclusion that they don't have PROOF of a problem to certify a hacked election....hmmm....political convention meets statistics!
The size of the probability is not as important as the conclusion and the actions we take in the future.
_________________________________________________________________________________
Sancho replied to OnTheOtherHand
You resort to rhetoric...
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=203x460141#460337
1.) Precinct level data would be very useful to generalize the samples in some polls, and compare to other precinct and district data from other election information. Exactly some of the criticism of TIA.
2.) Pre and Exit polls are generated and published by media...who relish in announcing the results. That means they should relish alternative analysis of the data published and be responsible for accuracy.
3.) Don't criticize TIA's between poll error (or WPE or anything else) when the pollsters can control this issue, but don't want to...at least they don't appear to try.
4.) Banning from DU is off the thread (as you like to say)
5.) You are now guilty of "selecting" the data to suit your argument - as R.A.
Fisher in the song I posted for you. TIA and EDA and others clearly state the polls used...and TIA tends to use everything available. We all know about the pre-election polls, trends, and predictions.
6.) Where have you been...EDA and others name the questions and details in their reports.
7.) You don't have to like the question, just answer it! Why don't pollsters investigate and focus on the interesting races and districts? They must not be interested!
8.) If you want to be critical, then defend your argument: if TIA doesn't meet "statistical assumptions", then how do you know? How do you know that TIA's analysis is not "robust" in terms of the "assumptions"? You are guessing just as you accuse TIA of starting with faulty assumption. I've seen little real evidence of either because of number 1 and 6 above!
As I stated to start with, there is NO amount of evidence that would convince you if you intend to defend pollsters or attack TIA..."those convinced against their will are of the same opinion still"
I'm still open to the possibility that TIA may not have the "perfect" formula or exactly correct "probability", but there is certainly some interesting merit to this latest argument that awaits discussion.
_________________________________________________________________________________
Another Sancho reply to OTOH:
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=203x460141#460340
Have you tested any of the pre (or exit polls) or election results for that matter for normality?
Given some common questions, have you tested that polls come from different populations?
My early analysis of precinct level data where available shows a common population so far...but mostly I've been looking at precincts of interest to me in Florida. If you want, I can start another thread with some Florida data already in SPSS or SYSTAT for those who are interested, but it would likely be pretty dull. I've also found some misfit in the categories in the questions in the 2004 E-M data...interesting, but that is not what you want.
The point is that you may be critical, but you don't really demonstrate that TIA is wrong any more than he demonstrates he is right based on unexplored/unavailable data...it's unknown. And there is no opportunity to obtain the "assumptions" that you say TIA are missing.
When I read articles and links referenced by Febble, I see some interesting things, and other things that do more manipulation than digging for the answers.
I don't want to debate power and effect size on DU, but that doesn't explain that there are lots of debates that can be settled. We would all like to see election officials do a better job, but one way to force the issue is to report if there is or is not a bunch of poll data that reveal an issue in a single disputed race or precinct or district that can't be explained in any way by "poll errors".
To do that, the pollsters (pre or exit) simply need to want to do it...and they don't want to. Are they chickens or false prophets or protecting clients or what? The answer that they don't want to know how people voted is simply not acceptable any more...and criticisms that follow from those polls that researchers don't meet assumptions (such as on this thread) are misleading if TIA or others do the best with the limited data they have...
IF TIA and EDA designed and advised the polling, I'd bet the quality of data would satisfy them one way or the other.
_________________________________________________________________________________
Sancho again:
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&
address=203x460141#460412
I didn't say TIA was correct...in fact, his probabilities are likely overstated...and I've said so before.
What I have advocated is that some the "assumptions" discussed are not testable given the data, and may not matter if the conclusion is robust to a violation of the mathematical assumption. That may be too technical for this thread. If you want to discuss specific objectivity, independence, sampling vs. sample distributions, levels of data quality, etc..we're getting a little difficult for the general audience. When Febble mentions within/between variance, that is ONE of the issues, but pollsters tend to play up what they are familiar with and ignore other things. Most of the "assumptions" could be addressed by good poll design and sampling plans and transparent data access. Some poll data is flawed to start with (like the assumption that likert scales are interval level). Some are collected in flawed ways to start with (like failure to sample the nonignorable nonrespondent).
It is not correct to criticize TIA for "not meeting assumptions" since there is no way to know if assumptions are met nor if it makes a difference in the conclusions that "polls don't match the election". That is a false assertion. Neither you nor TIA KNOW if all the "mathematical assumptions" are met!
There are other ways to skin the cat and Febble is correct that we may be well-served to look at things more visible and relevant.
Some of TIA (and EDA) discoveries are descriptively interesting and deserve investigation, regardless of the "magnitude" of the statistics.
Statisticians and engineers disagreed on test data that predicted the space shuttle would blow up in the cold because"it had not been empirically tested and there was not data that met mathematical rigor". Guess who was right!
Papers and pundits down here use the lack of "proof" in the polls (of manipulation) to AVOID fixing the election system, just like cigarette manufacturers used lack of "proof" of causes of cancer for decades due to "statistical significance". Regardless of the effect size, TIA and EDA force the issue to the surface.
I still think pollsters are irresponsible (unlikely) or incompetent (unlikely) or scared to piss off the paycheck (likely). There may be a combination of the three. Otherwise, we would see more action to address the issues.
Which congressional candidate are you going to vote for, the Democrat or Republican?
I created a trend line model based on a simple linear ("best-fit") regression of the 120 poll shares.
The model projected the Democrats would win 56.43% of the House Generic vote.
The Linear Regression Trend Model produced the following "best fit" trend lines.
Dem = 46.98 + .0419x
Rep = 38.06 + .0047x
Substituting x = 120 and allocating 60% of undecideds (UVA) to the Democrats:
..... Trend + UVA = Projection
Dem = 52.01 + 4.42 = 56.43%
Rep = 38.62 + 2.95 = 41.57%
1http://img.photobucket.com/albums/v474/autorank/Election2006_16921_image001.png
Exit pollsters Edison-Mitofsky asked 13,251 responders:
Who did you just vote for, the Democrat or Republican?
http://www.ropercenter.uconn.edu/elections/common/exitpolls.html
6351 respondents were asked who they voted for in 2004
The unadjusted exit poll of 13251 respondents indicated that the Democrats won the generic vote by:
Dem Rep
7470 5476
56.37% 41.33%
The pre-election Generic 120 poll trend matched the exit poll to within .07%!
Scenario 1:
Assume the unadjusted exit poll mix of returning Kerry and Bush voters:
Adjust vote shares to match the unadjusted exit poll.
Result:
The Final NEP was forced to match the 52.2-45.9% recorded margin by
1) adjusting the returning Kerry/Bush voter mix to an implausible 43/49%
2) lowering the Democratic shares of returning Kerry and Bush voters by approximately 2-3%.
In order to match the 56.37-41.33% vote split using 7pm NEP vote shares, the returning voter mix had to be revised:
1) the impossible 4.0% Other mix was lowered to 1.0% to match the 2004 recorded share.
2) the Kerry/Bush returning voter mix was changed to 48.6/46.4%.
Bush vote shares had to be lowered to match the unadjusted exit poll.
Adjusted 7pm NEP shares Final Final NEP shares
MIX Dem Rep Other MIX Dem Rep Other
Kerry 48.6% 93% 6% 1% 43% 92% 7% 1%
Bush 46.4% 17% 81% 2% 49% 15% 83% 2%
Other 1.0% 65% 18% 15% 4% 66% 23% 11%
DNV 4.0% 66% 16% 19% 4% 66% 32% 2%
TOTAL 100.00% 56.38% 41.32% 2.32% 100% 52.19% 45.88% 1.93%
Margin 15.06% 6.31%
This is further evidence that Kerry must have won in 2004.
______________________________________________________________________
The Final NEP radically changed the 7pm MIX and shares in order match the recorded vote.
Note that even the 7pm returning voter mix was highly dubious since it assumes that
the 2004 recorded vote was the True Vote (we know it wasn't)
7pm NEP Final NEP
MIX Dem Rep Other MIX Dem Rep Other
Kerry 45% 93% 6% 1% 43% 92% 7% 1%
Bush 47% 17% 82% 1% 49% 15% 83% 2%
Other 4% 66% 24% 10% 4% 66% 23% 11%
DNV 4% 67% 30% 3% 4% 66% 32% 2%
TOTAL 100% 55.16% 43.40% 1.44% 100% 52.19% 45.88% 1.93%
Margin 11.76% 6.31%
___________________________________________________________________________________
Scenario 2:
Assume that
1) The returning voter mix was proportional to the 2004 exit poll (Kerry 52-47%)
2) Returning Other voters were 1% of the 2006 total
3) Preliminary 7pm NEP vote shares
The resulting Democratic vote share is 57.87%.
In order to match the recorded vote, the Final NEP had to be adjusted by:
1) reducing the 7pm NEP Democratic vote shares
2) reversing the returning mix from +6% Kerry to +6% Bush
a) increasing the Other mix from 1% to 4% by deducting 3% from Kerry.
b) increasing the Bush mix by and decreasing the Kerry mix.
Unadjusted NEP Final NEP
MIX Dem Rep Other MIX Dem Rep Other
Kerry 50.5% 93% 6% 1% 43% 92% 7% 1%
Bush 44.5% 17% 82% 1% 49% 15% 83% 2%
Other 1% 66% 24% 10% 4% 66% 23% 11%
DNV 4% 67% 30% 3% 4% 66% 32% 2%
TOTAL 100% 57.87% 40.96% 1.17% 100% 52.19% 45.88% 1.93%
_____________________________________________________________________________
Now for a little history. In 2006, I posted the results of the 120 Generic Poll Trend model on Scoop, DU and PI (Dem 56.4 - Rep 41.6%).
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=364x2600324
But the Final NEP (Dem 52-Rep 46%) was way off (as usual, it was forced to match the recorded vote):
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=203x459299
So I posted an analysis which calculated a near zero probability that the recorded vote would deviate by over 4% from the 120 poll trendline.
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=364x2775205
Immediately, the usual DU naysayers jumped in. Surprisingly, Skinner joined them. He offered a lengthy polemic in which he lambasted my use of pre-election Generic polls to project the House vote. As far as I know, Skinner was not experienced in statistical polling analysis (he never commented in the pre-election and exit poll debates which permeated DU in 2004-2005. Apparently, he became an expert overnight. And now he was going to teach DUers all about Generic polls - and bash my "rookie" analysis in the process.
He claimed that pre-election Generic Polls should not be used for projecting the vote. He used caps for emphasis (I used to use caps a lot, too): It is "WRONG, WRONG, WRONG". He admitted that he didn't check my mathematics, but conjectured that I may have also made errors in the calculations.
He called my analysis an "embarrassment". He told DUers to THINK.
He ended with this advice: "You can do the best, most accurate, most awesome mathematics in the history of the world, but if you start with completely false assumptions, your "analysis" is going to be worthless".
Apparently, Skinner was unaware of the fact that for 40 years I have developed mathematically based scientific and financial models for many of the world's largest defense, consumer and financial corporations. Or that I developed and marketed corporate financial software to major banks and consumer product manufacturers. I assume Skinner is still in his thirties, so I was doing this before he was born.
_____________________________________________________________________
Despite the fact that the pre-election generic polls (adjusted for the 60% Democratic share of undecided voters) matched the 7pm national exit poll, Skinner, the new self-proclaimed polling "expert", wrote the following:
To those of you who keep demanding to see the problem, here it is:
http://www.democraticunderground.com/discuss/duboard.php?az=show_mesg&forum=364&topic_id=2775205&mesg_id=2781808
The problem is not in the mathematics (although I have not checked the math, so it's possible that there are errors there, too). The problem is in the assumptions he used before he even started.
TIA assumes that the "generic poll" should match the recorded votes. This is, quite simply, WRONG WRONG WRONG WRONG WRONG.
Does everyone here know what the "generic poll" is? The generic poll (usually referred to as the "generic congressional ballot") asks respondents which political party they support in the upcoming congressional election -- but it does not provide any names of any candidates. The generic congressional ballot is not -- and was never intended to be -- an accurate prediction of how people will vote. The point of the generic congressional ballot is to get a general sense of the mood of the voters.
Think, people. THINK.
Most American voters are not political junkies or activists. Even during an election season they could not tell you for certain the name of their member of Congress. When they get a call asking them their preference between a generic Democratic candidate and a generic Republican candidate, they will simply respond based on party. A large proportion of voters would be unable to think of the name of the candidate that represents each party in their upcoming congressional election.
But what happens when those same voters enter the voting booth on election day? They are presented with a list of NAMES OF REAL PEOPLE, their candidates, listed by office and political party. And they remember what they like and dislike about all of these people. They might have expressed support for a particular party, and they might hold that generic opinion of that party, but they like their own member of Congress and don't particularly care what party he or she is in. Or they might remember the ads about how a particular candidate was a crook or sexual deviant, and now that they finally face the realization he or she is running in their congressional district, they cannot possibly vote for him or her regardless of what party he represents.
Bottom line:
You can do the best, most accurate, most awesome mathematics in the history of the world, but if you start with completely false assumptions, your "analysis" is going to be worthless.
If you want to see how the outcome compares to the pre-election polls, you need to look at the pre-election polls that list candidates by name from each and every congressional district. Someone mentione up-thread that this is precisely what folks like Charlie Cook and Stu Rothenberg did before the election, and their predictions were quite accurate.
________________________________________________________________
I responded to that drivel with this post:
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=203x460141
I provided examples to show that the Generic polls ALL asked the proper question:
FOX News/Opinion Dynamics Poll. Nov. 4-5, 2006.
N=900 likely voters nationwide. MoE ± 3.
"Thinking ahead to this November's elections, if the congressional election were held today, would you vote for the Democratic candidate in your district or the Republican candidate in your district?" If unsure: "Well, if you had to vote, which way would you lean?"
I reviewed the Generic poll projection model:
The projected Democratic vote share was based on the trend line of ALL 120 Generic polls taken from Sept. 2005 up to Election Day. The Democrats won ALL 120 polls by an average 13.24% margin, 51.84D-38.60R, with 2% going to 3rd party candidates. I allocated 60% (UVA) of the 7.56% undecided voters to the Democrats. The final projection was 56.43D-41.62R, a 14.81% margin.
I explained the rationale for using a 1.5% MoE to compute the probability of the discrepancy:
One DUer made the following statement regarding the analysis: "Nowhere does it show how that polling sample would be 25% less variable than a standard poll, which almost always round DOWN (due to sample size) to 2%. That would inflate the distribution of the final analysis by a factor of 4/3 which moves the tails farther from the mean. Since the distributions are not linear, at that end of the curve, it could move the probabilities out by a factor of more than 100".
Here's why the 1.5% MoE was justified: I used the FINAL COMBINED 10 polls to calculate the MoE. There were 1000 sampled in each poll. Ten (10) INDEPENDENT Final Generic Polls are essentially equivalent to ONE poll of 10,000 sample size. The MoE for a 10,000 sample is near 1.0%. So the 1.5% MoE assumption was a conservative one.
The formula used to calculate MoE is:
MoE = 1.96*standard error = 1.96*SQRT((1-p)*p)/n)
For p=.56 and n= 10,000 sample-size, MoE = 0.97%.
I reviewed Generic polls and basic statistics:
Generic polls are designed to sample representative congressional districts. They all ask the same question. Calculating an average trend line or arithmetic mean gives us greater confidence that the sample mean is close to the true population mean. Is there anyone who will question the Law of Large Numbers and the Central Limit Theorem?
_________________________________________________________________________________
Skinner responded to the OP with a non-response. He obviously realized that he was just called out:
I'll save myself the effort and link to my previous post.
http://www.democraticunderground.com/discuss/duboard.php?az=show_mesg&forum=364&topic_id=2775205&mesg_id=2781808
Because my response is still correct, and TIA's central assumption is still wrong. He can try to dazzle people with a wall of words and numbers, but it will not matter. He cannot turn falsehood into truth simply because he wishes it to be so.
_________________________________________________________________________________
anaxarchos replied to Skinner:
You might want to exert that effort...
He proceeded to provide evidence that reputable pollsters and political scientists use Generic polls in their projections:
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=203x460141#460379
The central thesis in your link is summed up by this:
"If you want to see how the outcome compares to the pre-election polls, you need to look at the pre-election polls that list candidates by name from each and every congressional district. Someone mentioned up-thread that this is precisely what folks like Charlie Cook and Stu Rothenberg did before the election, and their predictions were quite accurate."
Even a casual search on Charlie Cook and Stu Rothenberg and "generic polls" yields dozens of both pre-election and post-election comments by both using generic polls in a similar manner to TIA. Certainly, they checked the generics against specific contests and other polling but, in this election, the generics seem to have been the most important instrument used by both and this, both in their detailed projections and in their post-mortems.
There have been problems with generics in the past, but, they were clearly useful in 2006 and neither Cook nor Rothenberg is a very good source for your blanket dismissal. Given that, you may want to explain why a compilation of over one hundred generics constitutes "completely false assumptions" or a "worthless" analysis.
http://www.cookpolitical.com /
November 6, 2006
All Monday there was considerable talk that the national picture had suddenly changed and that there was a significant tightening in the election. This was based in part on two national polls that showed the generic congressional ballot test having tightened to four (Pew) and six (ABC/Wash Post) points.
(snip)
Furthermore, there is no evidence of a trend in the generic ballot test. In chronological order of interviewing (using the midpoint of field dates), the margins were: 15 points (Time 11/1-3), 6 points (ABC/Wash Post), 4 points (Pew), 7 points (Gallup), 16 points (Newsweek), 20 points (CNN) and 13 points (Fox).
In individual races, some Republican pollsters see some movement, voters "coming home," in their direction, and/or some increase in intensity among GOP voters. All seem to think that it was too little, too late to significantly change the outcome. However, it might be enough to save a few candidates. None think it is a major change in the dynamics of races, and most remain somewhere between fairly and extremely pessimistic about tomorrow's outcome.
http://blog.washingtonpost.com/thefix/2006/08/parsing_t...
The answer, according to Charlie Cook and Stu Rothenberg, is a guarded yes.
"If you take an average of the last three or four polls, because any one can be an outlier in either direction, you can determine which way the wind is blowing, and whether the wind speed is small, medium, large or extra-large," said Cook. "The last three generics that I have seen have been in the 18 or 19 point range, which is on the high side of extra large. That suggests the probability of large Democratic gains."
"The generic surely reflects voters dissatisfaction with the President and his party and their inclination to support Democrats in the fall," agreed Rothenberg. "The size of the Democrats' generic advantage also can't be ignored. It too suggests the likelihood of a partisan wave, even though it does not guarantee the fate of any individual Republican incumbent."
http://rothenbergpoliticalreport.blogspot.com/2006/07/w...
Second, the ballot test in the July 10-13 Cooper & Secrest poll strongly mirrors the generic ballot in the district. Donnelly leads Chocola by 10 points in the ballot test (48 percent to 38 percent) while a generic Democratic candidate leads a generic Republican by 10 points as well (46 percent to 36 percent). The Democrats’ generic ballot advantage grew from 1 point in November 2005 to 10 points earlier this month, which also helps explain Donnelly’s improved standing in the July survey.
Personally, I'm not sure about the probability or extent of fraud in 2006... or in the practicality of mass fraud in mid-term congressional elections in general. Nevertheless, dismissing that possibility is not as easy as we might want it to be.
_________________________________________________________________________________
Skinner weakly replied to anaxarchos:
107. They're using the generic ballot in a similar manner to TIA? Oh really?
Find me the place where they use generics to argue that millions of Democratic votes were lost. Or find me a place where they claim the generic ballot provides an accurate prediction of the final vote count.
Do they use generics to inform their predictions? Of course they do. Generic ballots are a useful tool. I have never argued that they are not, and I have never argued that they are completely worthless. What I have said is that they are not supposed to accurately predict the final vote. All of the quotes you have provided in your post are fully consistent with my argument.
And, for the record, I am not "dissmissing the possibility of fraud." I consider it offensive and misleading to suggest that I have some sort of hidden nefarious agenda. I fully support any and all efforts to secure our elections and to prevent and uncover fraud. But that doesn't mean that I am required to shut off my brain and pretend not to notice when people on my own side make glaring errors. TIA may not care, but in the real world credibility matters. You can't keep making very basic rookie mistakes without people eventually dismissing the entire effort as a fabrication.
_________________________________________________________________________________
Sancho, a very well-respected FL Election Supervisor, commented:
167. A good point by TIA....from the scholars!
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&
address=203x460141#461262
As I've been learning about the world of pollsters (thanks to links from Febble and my own searching), I agree with much that I found in these articles from a single volumn of POQ. The exit pollsters have sacrificed accuracy for a single poll that is fast, but has too much error. TIA is appropriately using multiple sources that serve as checks and balances against each other, or multiple sources of evidence.
It seems to me, that many of the criticisms of the VNS exit polls are intuitively as serious as those committed by TIA, even if TIA's "math" and "assumptions" are not peer reviewed! Mitofsky (if you read between the lines) pretty much describes the poll failures (at that time) in "technical terms", and the connection of pre to exit is a suggested solution, as is the use of multiple polls, as is appropriate sample sizes, etc. In effect, VNS called elections with logic mistakes at least as "bad" as TIA is accused of on this blog, but the VNS had the responsibility and paycheck to get it right!
The pollsters also refuse to consider fraud (as alleged here on occasion), but admit they cut corners. The original poll designers also describe TV station analysis performing mathematical projections (comparison with previous year's polls) without the power and data to do the job, essentially not meeting the assumptions. Because non-response error is common and often cited doesn't mean that the more exotic, but equally bad mistakes by the VNS don't deserve attention; and considering fraud as a source of variation is clearly missing from the "scholars" as often as TIA suggests it as the problem.
_________________________________________________________________________________
Sancho replied to Febble:
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=203x460141#460182
TIA's formula "mistake" affects power as much as anything. Regardless, even the "odds" that 100 to 1, much less 100,000,000 to 1 that an election is hacked is pretty serious. I think that political correctness among pollsters is to avoid making the accusation, and I now accept that...but what can be done by pollsters to help with post-hoc analyses that would demand a different system or revote? Florida judges will USE the POLLSTER'S conclusion that they don't have PROOF of a problem to certify a hacked election....hmmm....political convention meets statistics!
The size of the probability is not as important as the conclusion and the actions we take in the future.
_________________________________________________________________________________
Sancho replied to OnTheOtherHand
You resort to rhetoric...
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=203x460141#460337
1.) Precinct level data would be very useful to generalize the samples in some polls, and compare to other precinct and district data from other election information. Exactly some of the criticism of TIA.
2.) Pre and Exit polls are generated and published by media...who relish in announcing the results. That means they should relish alternative analysis of the data published and be responsible for accuracy.
3.) Don't criticize TIA's between poll error (or WPE or anything else) when the pollsters can control this issue, but don't want to...at least they don't appear to try.
4.) Banning from DU is off the thread (as you like to say)
5.) You are now guilty of "selecting" the data to suit your argument - as R.A.
Fisher in the song I posted for you. TIA and EDA and others clearly state the polls used...and TIA tends to use everything available. We all know about the pre-election polls, trends, and predictions.
6.) Where have you been...EDA and others name the questions and details in their reports.
7.) You don't have to like the question, just answer it! Why don't pollsters investigate and focus on the interesting races and districts? They must not be interested!
8.) If you want to be critical, then defend your argument: if TIA doesn't meet "statistical assumptions", then how do you know? How do you know that TIA's analysis is not "robust" in terms of the "assumptions"? You are guessing just as you accuse TIA of starting with faulty assumption. I've seen little real evidence of either because of number 1 and 6 above!
As I stated to start with, there is NO amount of evidence that would convince you if you intend to defend pollsters or attack TIA..."those convinced against their will are of the same opinion still"
I'm still open to the possibility that TIA may not have the "perfect" formula or exactly correct "probability", but there is certainly some interesting merit to this latest argument that awaits discussion.
_________________________________________________________________________________
Another Sancho reply to OTOH:
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=203x460141#460340
Have you tested any of the pre (or exit polls) or election results for that matter for normality?
Given some common questions, have you tested that polls come from different populations?
My early analysis of precinct level data where available shows a common population so far...but mostly I've been looking at precincts of interest to me in Florida. If you want, I can start another thread with some Florida data already in SPSS or SYSTAT for those who are interested, but it would likely be pretty dull. I've also found some misfit in the categories in the questions in the 2004 E-M data...interesting, but that is not what you want.
The point is that you may be critical, but you don't really demonstrate that TIA is wrong any more than he demonstrates he is right based on unexplored/unavailable data...it's unknown. And there is no opportunity to obtain the "assumptions" that you say TIA are missing.
When I read articles and links referenced by Febble, I see some interesting things, and other things that do more manipulation than digging for the answers.
I don't want to debate power and effect size on DU, but that doesn't explain that there are lots of debates that can be settled. We would all like to see election officials do a better job, but one way to force the issue is to report if there is or is not a bunch of poll data that reveal an issue in a single disputed race or precinct or district that can't be explained in any way by "poll errors".
To do that, the pollsters (pre or exit) simply need to want to do it...and they don't want to. Are they chickens or false prophets or protecting clients or what? The answer that they don't want to know how people voted is simply not acceptable any more...and criticisms that follow from those polls that researchers don't meet assumptions (such as on this thread) are misleading if TIA or others do the best with the limited data they have...
IF TIA and EDA designed and advised the polling, I'd bet the quality of data would satisfy them one way or the other.
_________________________________________________________________________________
Sancho again:
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&
address=203x460141#460412
I didn't say TIA was correct...in fact, his probabilities are likely overstated...and I've said so before.
What I have advocated is that some the "assumptions" discussed are not testable given the data, and may not matter if the conclusion is robust to a violation of the mathematical assumption. That may be too technical for this thread. If you want to discuss specific objectivity, independence, sampling vs. sample distributions, levels of data quality, etc..we're getting a little difficult for the general audience. When Febble mentions within/between variance, that is ONE of the issues, but pollsters tend to play up what they are familiar with and ignore other things. Most of the "assumptions" could be addressed by good poll design and sampling plans and transparent data access. Some poll data is flawed to start with (like the assumption that likert scales are interval level). Some are collected in flawed ways to start with (like failure to sample the nonignorable nonrespondent).
It is not correct to criticize TIA for "not meeting assumptions" since there is no way to know if assumptions are met nor if it makes a difference in the conclusions that "polls don't match the election". That is a false assertion. Neither you nor TIA KNOW if all the "mathematical assumptions" are met!
There are other ways to skin the cat and Febble is correct that we may be well-served to look at things more visible and relevant.
Some of TIA (and EDA) discoveries are descriptively interesting and deserve investigation, regardless of the "magnitude" of the statistics.
Statisticians and engineers disagreed on test data that predicted the space shuttle would blow up in the cold because"it had not been empirically tested and there was not data that met mathematical rigor". Guess who was right!
Papers and pundits down here use the lack of "proof" in the polls (of manipulation) to AVOID fixing the election system, just like cigarette manufacturers used lack of "proof" of causes of cancer for decades due to "statistical significance". Regardless of the effect size, TIA and EDA force the issue to the surface.
I still think pollsters are irresponsible (unlikely) or incompetent (unlikely) or scared to piss off the paycheck (likely). There may be a combination of the three. Otherwise, we would see more action to address the issues.