Log in

View Full Version : 2006 midterms: 120 pre-election Generic poll trend (Dem 56.43- Rep 41.57%) matched the unadjusted NEP (56.37-41.33%)



TruthIsAll
03-26-2009, 10:50 AM
In 2006, 120 Pre-election House Generic Polls asked the question:
Which congressional candidate are you going to vote for, the Democrat or Republican?

I created a trend line model based on a simple linear ("best-fit") regression of the 120 poll shares.
The model projected the Democrats would win 56.43% of the House Generic vote.

The Linear Regression Trend Model produced the following "best fit" trend lines.
Dem = 46.98 + .0419x
Rep = 38.06 + .0047x

Substituting x = 120 and allocating 60% of undecideds (UVA) to the Democrats:
..... Trend + UVA = Projection
Dem = 52.01 + 4.42 = 56.43%
Rep = 38.62 + 2.95 = 41.57%

1http://img.photobucket.com/albums/v474/autorank/Election2006_16921_image001.png

Exit pollsters Edison-Mitofsky asked 13,251 responders:
Who did you just vote for, the Democrat or Republican?
http://www.ropercenter.uconn.edu/elections/common/exitpolls.html

6351 respondents were asked who they voted for in 2004

The unadjusted exit poll of 13251 respondents indicated that the Democrats won the generic vote by:
Dem Rep
7470 5476
56.37% 41.33%

The pre-election Generic 120 poll trend matched the exit poll to within .07%!

Scenario 1:
Assume the unadjusted exit poll mix of returning Kerry and Bush voters:
Adjust vote shares to match the unadjusted exit poll.

Result:
The Final NEP was forced to match the 52.2-45.9% recorded margin by
1) adjusting the returning Kerry/Bush voter mix to an implausible 43/49%
2) lowering the Democratic shares of returning Kerry and Bush voters by approximately 2-3%.




In order to match the 56.37-41.33% vote split using 7pm NEP vote shares, the returning voter mix had to be revised:

1) the impossible 4.0% Other mix was lowered to 1.0% to match the 2004 recorded share.
2) the Kerry/Bush returning voter mix was changed to 48.6/46.4%.
Bush vote shares had to be lowered to match the unadjusted exit poll.


Adjusted 7pm NEP shares Final Final NEP shares
MIX Dem Rep Other MIX Dem Rep Other
Kerry 48.6% 93% 6% 1% 43% 92% 7% 1%
Bush 46.4% 17% 81% 2% 49% 15% 83% 2%
Other 1.0% 65% 18% 15% 4% 66% 23% 11%
DNV 4.0% 66% 16% 19% 4% 66% 32% 2%

TOTAL 100.00% 56.38% 41.32% 2.32% 100% 52.19% 45.88% 1.93%
Margin 15.06% 6.31%


This is further evidence that Kerry must have won in 2004.
______________________________________________________________________


The Final NEP radically changed the 7pm MIX and shares in order match the recorded vote.
Note that even the 7pm returning voter mix was highly dubious since it assumes that
the 2004 recorded vote was the True Vote (we know it wasn't)

7pm NEP Final NEP
MIX Dem Rep Other MIX Dem Rep Other
Kerry 45% 93% 6% 1% 43% 92% 7% 1%
Bush 47% 17% 82% 1% 49% 15% 83% 2%
Other 4% 66% 24% 10% 4% 66% 23% 11%
DNV 4% 67% 30% 3% 4% 66% 32% 2%

TOTAL 100% 55.16% 43.40% 1.44% 100% 52.19% 45.88% 1.93%
Margin 11.76% 6.31%

___________________________________________________________________________________

Scenario 2:

Assume that
1) The returning voter mix was proportional to the 2004 exit poll (Kerry 52-47%)
2) Returning Other voters were 1% of the 2006 total
3) Preliminary 7pm NEP vote shares

The resulting Democratic vote share is 57.87%.

In order to match the recorded vote, the Final NEP had to be adjusted by:
1) reducing the 7pm NEP Democratic vote shares
2) reversing the returning mix from +6% Kerry to +6% Bush
a) increasing the Other mix from 1% to 4% by deducting 3% from Kerry.
b) increasing the Bush mix by and decreasing the Kerry mix.


Unadjusted NEP Final NEP
MIX Dem Rep Other MIX Dem Rep Other
Kerry 50.5% 93% 6% 1% 43% 92% 7% 1%
Bush 44.5% 17% 82% 1% 49% 15% 83% 2%
Other 1% 66% 24% 10% 4% 66% 23% 11%
DNV 4% 67% 30% 3% 4% 66% 32% 2%

TOTAL 100% 57.87% 40.96% 1.17% 100% 52.19% 45.88% 1.93%



_____________________________________________________________________________

Now for a little history. In 2006, I posted the results of the 120 Generic Poll Trend model on Scoop, DU and PI (Dem 56.4 - Rep 41.6%).
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=364x2600324

But the Final NEP (Dem 52-Rep 46%) was way off (as usual, it was forced to match the recorded vote):
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=203x459299

So I posted an analysis which calculated a near zero probability that the recorded vote would deviate by over 4% from the 120 poll trendline.
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=364x2775205

Immediately, the usual DU naysayers jumped in. Surprisingly, Skinner joined them. He offered a lengthy polemic in which he lambasted my use of pre-election Generic polls to project the House vote. As far as I know, Skinner was not experienced in statistical polling analysis (he never commented in the pre-election and exit poll debates which permeated DU in 2004-2005. Apparently, he became an expert overnight. And now he was going to teach DUers all about Generic polls - and bash my "rookie" analysis in the process.

He claimed that pre-election Generic Polls should not be used for projecting the vote. He used caps for emphasis (I used to use caps a lot, too): It is "WRONG, WRONG, WRONG". He admitted that he didn't check my mathematics, but conjectured that I may have also made errors in the calculations.

He called my analysis an "embarrassment". He told DUers to THINK.

He ended with this advice: "You can do the best, most accurate, most awesome mathematics in the history of the world, but if you start with completely false assumptions, your "analysis" is going to be worthless".

Apparently, Skinner was unaware of the fact that for 40 years I have developed mathematically based scientific and financial models for many of the world's largest defense, consumer and financial corporations. Or that I developed and marketed corporate financial software to major banks and consumer product manufacturers. I assume Skinner is still in his thirties, so I was doing this before he was born.

_____________________________________________________________________

Despite the fact that the pre-election generic polls (adjusted for the 60% Democratic share of undecided voters) matched the 7pm national exit poll, Skinner, the new self-proclaimed polling "expert", wrote the following:

To those of you who keep demanding to see the problem, here it is:
http://www.democraticunderground.com/discuss/duboard.php?az=show_mesg&forum=364&topic_id=2775205&mesg_id=2781808

The problem is not in the mathematics (although I have not checked the math, so it's possible that there are errors there, too). The problem is in the assumptions he used before he even started.

TIA assumes that the "generic poll" should match the recorded votes. This is, quite simply, WRONG WRONG WRONG WRONG WRONG.

Does everyone here know what the "generic poll" is? The generic poll (usually referred to as the "generic congressional ballot") asks respondents which political party they support in the upcoming congressional election -- but it does not provide any names of any candidates. The generic congressional ballot is not -- and was never intended to be -- an accurate prediction of how people will vote. The point of the generic congressional ballot is to get a general sense of the mood of the voters.

Think, people. THINK.

Most American voters are not political junkies or activists. Even during an election season they could not tell you for certain the name of their member of Congress. When they get a call asking them their preference between a generic Democratic candidate and a generic Republican candidate, they will simply respond based on party. A large proportion of voters would be unable to think of the name of the candidate that represents each party in their upcoming congressional election.

But what happens when those same voters enter the voting booth on election day? They are presented with a list of NAMES OF REAL PEOPLE, their candidates, listed by office and political party. And they remember what they like and dislike about all of these people. They might have expressed support for a particular party, and they might hold that generic opinion of that party, but they like their own member of Congress and don't particularly care what party he or she is in. Or they might remember the ads about how a particular candidate was a crook or sexual deviant, and now that they finally face the realization he or she is running in their congressional district, they cannot possibly vote for him or her regardless of what party he represents.

Bottom line:
You can do the best, most accurate, most awesome mathematics in the history of the world, but if you start with completely false assumptions, your "analysis" is going to be worthless.

If you want to see how the outcome compares to the pre-election polls, you need to look at the pre-election polls that list candidates by name from each and every congressional district. Someone mentione up-thread that this is precisely what folks like Charlie Cook and Stu Rothenberg did before the election, and their predictions were quite accurate.

________________________________________________________________

I responded to that drivel with this post:
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=203x460141

I provided examples to show that the Generic polls ALL asked the proper question:

FOX News/Opinion Dynamics Poll. Nov. 4-5, 2006.
N=900 likely voters nationwide. MoE ± 3.
"Thinking ahead to this November's elections, if the congressional election were held today, would you vote for the Democratic candidate in your district or the Republican candidate in your district?" If unsure: "Well, if you had to vote, which way would you lean?"


I reviewed the Generic poll projection model:

The projected Democratic vote share was based on the trend line of ALL 120 Generic polls taken from Sept. 2005 up to Election Day. The Democrats won ALL 120 polls by an average 13.24% margin, 51.84D-38.60R, with 2% going to 3rd party candidates. I allocated 60% (UVA) of the 7.56% undecided voters to the Democrats. The final projection was 56.43D-41.62R, a 14.81% margin.

I explained the rationale for using a 1.5% MoE to compute the probability of the discrepancy:

One DUer made the following statement regarding the analysis: "Nowhere does it show how that polling sample would be 25% less variable than a standard poll, which almost always round DOWN (due to sample size) to 2%. That would inflate the distribution of the final analysis by a factor of 4/3 which moves the tails farther from the mean. Since the distributions are not linear, at that end of the curve, it could move the probabilities out by a factor of more than 100".

Here's why the 1.5% MoE was justified: I used the FINAL COMBINED 10 polls to calculate the MoE. There were 1000 sampled in each poll. Ten (10) INDEPENDENT Final Generic Polls are essentially equivalent to ONE poll of 10,000 sample size. The MoE for a 10,000 sample is near 1.0%. So the 1.5% MoE assumption was a conservative one.

The formula used to calculate MoE is:
MoE = 1.96*standard error = 1.96*SQRT((1-p)*p)/n)
For p=.56 and n= 10,000 sample-size, MoE = 0.97%.

I reviewed Generic polls and basic statistics:
Generic polls are designed to sample representative congressional districts. They all ask the same question. Calculating an average trend line or arithmetic mean gives us greater confidence that the sample mean is close to the true population mean. Is there anyone who will question the Law of Large Numbers and the Central Limit Theorem?

_________________________________________________________________________________

Skinner responded to the OP with a non-response. He obviously realized that he was just called out:

I'll save myself the effort and link to my previous post.

http://www.democraticunderground.com/discuss/duboard.php?az=show_mesg&forum=364&topic_id=2775205&mesg_id=2781808

Because my response is still correct, and TIA's central assumption is still wrong. He can try to dazzle people with a wall of words and numbers, but it will not matter. He cannot turn falsehood into truth simply because he wishes it to be so.

_________________________________________________________________________________

anaxarchos replied to Skinner:

You might want to exert that effort...

He proceeded to provide evidence that reputable pollsters and political scientists use Generic polls in their projections:

http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=203x460141#460379

The central thesis in your link is summed up by this:
"If you want to see how the outcome compares to the pre-election polls, you need to look at the pre-election polls that list candidates by name from each and every congressional district. Someone mentioned up-thread that this is precisely what folks like Charlie Cook and Stu Rothenberg did before the election, and their predictions were quite accurate."

Even a casual search on Charlie Cook and Stu Rothenberg and "generic polls" yields dozens of both pre-election and post-election comments by both using generic polls in a similar manner to TIA. Certainly, they checked the generics against specific contests and other polling but, in this election, the generics seem to have been the most important instrument used by both and this, both in their detailed projections and in their post-mortems.

There have been problems with generics in the past, but, they were clearly useful in 2006 and neither Cook nor Rothenberg is a very good source for your blanket dismissal. Given that, you may want to explain why a compilation of over one hundred generics constitutes "completely false assumptions" or a "worthless" analysis.

http://www.cookpolitical.com /
November 6, 2006

All Monday there was considerable talk that the national picture had suddenly changed and that there was a significant tightening in the election. This was based in part on two national polls that showed the generic congressional ballot test having tightened to four (Pew) and six (ABC/Wash Post) points.

(snip)

Furthermore, there is no evidence of a trend in the generic ballot test. In chronological order of interviewing (using the midpoint of field dates), the margins were: 15 points (Time 11/1-3), 6 points (ABC/Wash Post), 4 points (Pew), 7 points (Gallup), 16 points (Newsweek), 20 points (CNN) and 13 points (Fox).

In individual races, some Republican pollsters see some movement, voters "coming home," in their direction, and/or some increase in intensity among GOP voters. All seem to think that it was too little, too late to significantly change the outcome. However, it might be enough to save a few candidates. None think it is a major change in the dynamics of races, and most remain somewhere between fairly and extremely pessimistic about tomorrow's outcome.

http://blog.washingtonpost.com/thefix/2006/08/parsing_t...

The answer, according to Charlie Cook and Stu Rothenberg, is a guarded yes.
"If you take an average of the last three or four polls, because any one can be an outlier in either direction, you can determine which way the wind is blowing, and whether the wind speed is small, medium, large or extra-large," said Cook. "The last three generics that I have seen have been in the 18 or 19 point range, which is on the high side of extra large. That suggests the probability of large Democratic gains."

"The generic surely reflects voters dissatisfaction with the President and his party and their inclination to support Democrats in the fall," agreed Rothenberg. "The size of the Democrats' generic advantage also can't be ignored. It too suggests the likelihood of a partisan wave, even though it does not guarantee the fate of any individual Republican incumbent."

http://rothenbergpoliticalreport.blogspot.com/2006/07/w...

Second, the ballot test in the July 10-13 Cooper & Secrest poll strongly mirrors the generic ballot in the district. Donnelly leads Chocola by 10 points in the ballot test (48 percent to 38 percent) while a generic Democratic candidate leads a generic Republican by 10 points as well (46 percent to 36 percent). The Democrats’ generic ballot advantage grew from 1 point in November 2005 to 10 points earlier this month, which also helps explain Donnelly’s improved standing in the July survey.

Personally, I'm not sure about the probability or extent of fraud in 2006... or in the practicality of mass fraud in mid-term congressional elections in general. Nevertheless, dismissing that possibility is not as easy as we might want it to be.

_________________________________________________________________________________

Skinner weakly replied to anaxarchos:
107. They're using the generic ballot in a similar manner to TIA? Oh really?

Find me the place where they use generics to argue that millions of Democratic votes were lost. Or find me a place where they claim the generic ballot provides an accurate prediction of the final vote count.

Do they use generics to inform their predictions? Of course they do. Generic ballots are a useful tool. I have never argued that they are not, and I have never argued that they are completely worthless. What I have said is that they are not supposed to accurately predict the final vote. All of the quotes you have provided in your post are fully consistent with my argument.

And, for the record, I am not "dissmissing the possibility of fraud." I consider it offensive and misleading to suggest that I have some sort of hidden nefarious agenda. I fully support any and all efforts to secure our elections and to prevent and uncover fraud. But that doesn't mean that I am required to shut off my brain and pretend not to notice when people on my own side make glaring errors. TIA may not care, but in the real world credibility matters. You can't keep making very basic rookie mistakes without people eventually dismissing the entire effort as a fabrication.
_________________________________________________________________________________

Sancho, a very well-respected FL Election Supervisor, commented:
167. A good point by TIA....from the scholars!
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&
address=203x460141#461262

As I've been learning about the world of pollsters (thanks to links from Febble and my own searching), I agree with much that I found in these articles from a single volumn of POQ. The exit pollsters have sacrificed accuracy for a single poll that is fast, but has too much error. TIA is appropriately using multiple sources that serve as checks and balances against each other, or multiple sources of evidence.

It seems to me, that many of the criticisms of the VNS exit polls are intuitively as serious as those committed by TIA, even if TIA's "math" and "assumptions" are not peer reviewed! Mitofsky (if you read between the lines) pretty much describes the poll failures (at that time) in "technical terms", and the connection of pre to exit is a suggested solution, as is the use of multiple polls, as is appropriate sample sizes, etc. In effect, VNS called elections with logic mistakes at least as "bad" as TIA is accused of on this blog, but the VNS had the responsibility and paycheck to get it right!

The pollsters also refuse to consider fraud (as alleged here on occasion), but admit they cut corners. The original poll designers also describe TV station analysis performing mathematical projections (comparison with previous year's polls) without the power and data to do the job, essentially not meeting the assumptions. Because non-response error is common and often cited doesn't mean that the more exotic, but equally bad mistakes by the VNS don't deserve attention; and considering fraud as a source of variation is clearly missing from the "scholars" as often as TIA suggests it as the problem.

_________________________________________________________________________________

Sancho replied to Febble:
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=203x460141#460182

TIA's formula "mistake" affects power as much as anything. Regardless, even the "odds" that 100 to 1, much less 100,000,000 to 1 that an election is hacked is pretty serious. I think that political correctness among pollsters is to avoid making the accusation, and I now accept that...but what can be done by pollsters to help with post-hoc analyses that would demand a different system or revote? Florida judges will USE the POLLSTER'S conclusion that they don't have PROOF of a problem to certify a hacked election....hmmm....political convention meets statistics!

The size of the probability is not as important as the conclusion and the actions we take in the future.

_________________________________________________________________________________

Sancho replied to OnTheOtherHand
You resort to rhetoric...

http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=203x460141#460337

1.) Precinct level data would be very useful to generalize the samples in some polls, and compare to other precinct and district data from other election information. Exactly some of the criticism of TIA.

2.) Pre and Exit polls are generated and published by media...who relish in announcing the results. That means they should relish alternative analysis of the data published and be responsible for accuracy.

3.) Don't criticize TIA's between poll error (or WPE or anything else) when the pollsters can control this issue, but don't want to...at least they don't appear to try.

4.) Banning from DU is off the thread (as you like to say)

5.) You are now guilty of "selecting" the data to suit your argument - as R.A.
Fisher in the song I posted for you. TIA and EDA and others clearly state the polls used...and TIA tends to use everything available. We all know about the pre-election polls, trends, and predictions.

6.) Where have you been...EDA and others name the questions and details in their reports.

7.) You don't have to like the question, just answer it! Why don't pollsters investigate and focus on the interesting races and districts? They must not be interested!

8.) If you want to be critical, then defend your argument: if TIA doesn't meet "statistical assumptions", then how do you know? How do you know that TIA's analysis is not "robust" in terms of the "assumptions"? You are guessing just as you accuse TIA of starting with faulty assumption. I've seen little real evidence of either because of number 1 and 6 above!

As I stated to start with, there is NO amount of evidence that would convince you if you intend to defend pollsters or attack TIA..."those convinced against their will are of the same opinion still"

I'm still open to the possibility that TIA may not have the "perfect" formula or exactly correct "probability", but there is certainly some interesting merit to this latest argument that awaits discussion.
_________________________________________________________________________________

Another Sancho reply to OTOH:
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&address=203x460141#460340

Have you tested any of the pre (or exit polls) or election results for that matter for normality?

Given some common questions, have you tested that polls come from different populations?

My early analysis of precinct level data where available shows a common population so far...but mostly I've been looking at precincts of interest to me in Florida. If you want, I can start another thread with some Florida data already in SPSS or SYSTAT for those who are interested, but it would likely be pretty dull. I've also found some misfit in the categories in the questions in the 2004 E-M data...interesting, but that is not what you want.

The point is that you may be critical, but you don't really demonstrate that TIA is wrong any more than he demonstrates he is right based on unexplored/unavailable data...it's unknown. And there is no opportunity to obtain the "assumptions" that you say TIA are missing.

When I read articles and links referenced by Febble, I see some interesting things, and other things that do more manipulation than digging for the answers.

I don't want to debate power and effect size on DU, but that doesn't explain that there are lots of debates that can be settled. We would all like to see election officials do a better job, but one way to force the issue is to report if there is or is not a bunch of poll data that reveal an issue in a single disputed race or precinct or district that can't be explained in any way by "poll errors".

To do that, the pollsters (pre or exit) simply need to want to do it...and they don't want to. Are they chickens or false prophets or protecting clients or what? The answer that they don't want to know how people voted is simply not acceptable any more...and criticisms that follow from those polls that researchers don't meet assumptions (such as on this thread) are misleading if TIA or others do the best with the limited data they have...

IF TIA and EDA designed and advised the polling, I'd bet the quality of data would satisfy them one way or the other.

_________________________________________________________________________________

Sancho again:
http://www.democraticunderground.com/discuss/duboard.php?az=view_all&
address=203x460141#460412

I didn't say TIA was correct...in fact, his probabilities are likely overstated...and I've said so before.

What I have advocated is that some the "assumptions" discussed are not testable given the data, and may not matter if the conclusion is robust to a violation of the mathematical assumption. That may be too technical for this thread. If you want to discuss specific objectivity, independence, sampling vs. sample distributions, levels of data quality, etc..we're getting a little difficult for the general audience. When Febble mentions within/between variance, that is ONE of the issues, but pollsters tend to play up what they are familiar with and ignore other things. Most of the "assumptions" could be addressed by good poll design and sampling plans and transparent data access. Some poll data is flawed to start with (like the assumption that likert scales are interval level). Some are collected in flawed ways to start with (like failure to sample the nonignorable nonrespondent).

It is not correct to criticize TIA for "not meeting assumptions" since there is no way to know if assumptions are met nor if it makes a difference in the conclusions that "polls don't match the election". That is a false assertion. Neither you nor TIA KNOW if all the "mathematical assumptions" are met!

There are other ways to skin the cat and Febble is correct that we may be well-served to look at things more visible and relevant.

Some of TIA (and EDA) discoveries are descriptively interesting and deserve investigation, regardless of the "magnitude" of the statistics.

Statisticians and engineers disagreed on test data that predicted the space shuttle would blow up in the cold because"it had not been empirically tested and there was not data that met mathematical rigor". Guess who was right!

Papers and pundits down here use the lack of "proof" in the polls (of manipulation) to AVOID fixing the election system, just like cigarette manufacturers used lack of "proof" of causes of cancer for decades due to "statistical significance". Regardless of the effect size, TIA and EDA force the issue to the surface.

I still think pollsters are irresponsible (unlikely) or incompetent (unlikely) or scared to piss off the paycheck (likely). There may be a combination of the three. Otherwise, we would see more action to address the issues.

TruthIsAll
04-06-2009, 07:02 PM
http://www.democraticunderground.com/discuss/duboard.php?az=show_topic&forum=203&topic_id=459299

The naysayers never quit, even when the 2006 election fraud analysis is vindicated by the numbers.


Febble (1000+ posts) Fri Nov-17-06 07:33 PM
Response to Reply #2
4. So, can you put the post into English?

It seems to me that what TIA is saying is that he predicted a bigger win than the Dems got, so there must have been fraud.

Do you know why TIA's predicted so much larger a gain for the Dems than other pundits, such as Larry Sabato?

http://www.centerforpolitics.org/crystalball /

For a while this forum looked as though it was focussed on the real evidence for corruption and miscounting in the last election.

I really hope that focus comes back. But haring off after frankly silly inferences made from polls is not doing anything for the credibility of the ER movement. You have a great case. You have power in the House and Senate. For God's sake don't blow it with crappy statistics.

. Statistics are crappy as a whole, or just TIAs statistics are crappy

I just want to make sure I understand you. K and R


Febble (1000+ posts) Sat Nov-18-06 04:53 AM
Response to Reply #6
18. Certainly not all statistical analyses

are crappy. There is some excellent work being done undervotes in Florida as we speak. But any statistical analysis is only as good as its assumptions, and TIAs are questionable. You need to put probabilities on your assumptions being correct, as well as on the probability of a particular inference being correct, given your assumptions. TIA doesn't.
Printer Friendly | Permalink | | Top

kster (1000+ posts) Sun Nov-19-06 02:18 AM
Response to Reply #18
42. Are the statistical analysis in your Country crappy...nt



Febble (1000+ posts) Sun Nov-19-06 08:12 AM
Response to Reply #42
45. Some are, some aren't

bad statistics are certainly not a uniquely US phenomenon. I have a paper on my desk full of bad statistics waiting for my review right now. I think the authors are Australian.

On the other hand some of the best statisticians I have ever met have been American including those who taught me.
Printer Friendly | Permalink | | Top

philb (1000+ posts) Sun Nov-19-06 11:10 PM
Response to Reply #18
50. That can be assessed to some degree independently of his analysis

If an argument is valid and the hypothesis is true, then the conclusion is true.

So the issue seems to be to define the argument and hypothesis, and assess whether the argument is valid and the hypothesis(or hypotheses) true. Or as you note the probability that the hypothesis is true.

And if the argument isn't clearly always valid, the probability that it is valid.


What do you think about the probability of his hypotheses being true?



Febble (1000+ posts) Mon Nov-20-06 03:38 AM
Response to Reply #50
52. Well, probabilities

Edited on Mon Nov-20-06 03:40 AM by Febble
regarding his assumptions would be somewhat subjective, but to estimate the likelihood of them thoroughly, you would have to look at past data. It's what pundits and political and social scientists, and public opinion survey researchers spend many hours, many conferences, and many papers doing. And in neither 2004 or 2006 was the consensus of opinion close to TIA's estimates. Sam Wang was briefly in line with TIA just before 2004, but he's not actually a political scientist, and he had the grace to make the point that he hadn't factored in the probability of his assumptions being true.

That in itself doesn't mean that TIA wrong, but it does mean he's out on a limb. And he seems to think that either no-one else has looked at the data, or they have and they are all out of step except him*.

If you are really interested in a detailed critique of TIA's analysis, I can take specifics. But his questionable assumptions include:

That sampling error is the only error in a poll.
That when there IS bias in the poll it would favor Republicans, not Democrats
That people recall their previous votes correctly (there is excellent evidence that they do not, and that they tend to misreport having voted for the previous winner)
That undecideds always break for the challenger
That incumbents below 50% always lose
That Likely Voter models are less reliable than Registered Models.

He also seems to misunderstand a lot about the way the exit poll data is collected and subsequently processed. I posted a diary about exit poll methodology on Daily Kos just before the election.

http://www.dailykos.com/storyonly/2006/11/4/135126/905

It would be nice to see evidence that TIA had actually read it, or something like it.




mom cat (1000+ posts) Mon Nov-20-06 06:58 PM
Response to Reply #52
55. TIA RESPONDS: Setting the record straight

As is typical, we can always expect beautifully stated misrepresentations of the facts as well as factual omissions in your posts. This one was no exception. Now I will set the record straight.

You claim that I made the following statements:

FEBBLE:
That sampling error is the only error in a poll.

TIA:
I never said that. I said that scientific polling minimizes the error and that’s why pollsters always quote the MoE based on sample-size.

FEBBLE:
That when there IS bias in the poll it would favor Republicans, not Democrats

TIA:
A few FACTS.
1) It’s a well known fact that approximately 3% of the votes are uncounted due to lost, spoiled and provisional ballots. The vast majority of these votes are in minority districts(50% in black districts). Since the minority black vote is
90% democratic, and the Hispanic vote at least 60%, a fair estimate is that 75% of the uncounted votes (2.25%) are Democrat and the other 0.75% Republican. That’s a 1.5% bias in favor of the Republicans.

2) There is a proven reluctance of low-income Democratic voters (under $50K) and high-income Republicans ($100k+) to participate in exit polls. But since low income voters outnumber high-income voters by almost 3-1, there's another component of Republican bias.

3) Exit poll response was high in Bush states and low in Kerry states.
Here's GRAPHIC MATHEMATICAL PROOF USING LINEAR REGRESSION ANALYSIS:

1http://www.geocities.com/electionmodel/StateVotevsExitPollCompletionRate1_27680_image001.png

FEBBLE:
That people recall their previous votes correctly (there is excellent evidence that they do not, and that they tend to misreport having voted for the previous winner)

TIA:
If you consider that the FINAL National Exit poll is ALWAYS MATCHED TO THE RECORDED VOTE, WHICH MEANS THAT THE WEIGHTS ARE FUDGED TO MATCH THE VOTE. I CALL IT FUDGING; YOU CALL IT VOTER MISREPRESENTATION. Witness the 2004 Final NEP, which Bush won by 51-48%: 43% of 2004 voters were Bush voters and 37% were Gore voters? This was a FUDGE necessary to match the recorded vote. In the 12:22am poll, which Kerry won by 51-48%, the mix was 41 Bush/39 Gore. Mathematically, the Bush MAXIMUM Bush 2000 representation weighting was 39.8%, which is the ratio of Bush 2000 voters still alive n 2004 by the 122.3mm who voted.

Do you want MORE of this evidence? Look at the 2006 NEP. The 7pm poll had the 2004 weighting as 45Bush/46 Gore. The FINAL had it 49 Bush/43 Gore/8 Other. Where did the 8% for OTHER voters come from? Third-party 2004 voters comprised 1% of the vote. Where did the excess 7% come from? Kerry. Here’s why:

Are we to believe Bush voters outnumbered Kerry voters by 6%? Even if you believe the 2004 FINAL, which we have proven bogus, the Bush margin was 3%. In reality it should have been 50 Kerry-47 Bush, after deducting 1% for voter mortality over the 2-year period. Now 50% = 43% + 7%. There is your Kerry vote.

FEBBLE:
That undecideds always break for the challenger.

TIA:
I never said ALWAYS. I said the vast majority of the time. The evidence is a study of 155 incumbent elections that you yourself quoted: in 82%, the challenger won the undecided vote; the incumbent, just 12%; neither, 6%. And of course, pollsters Zogby, Harriss and others have always claimed that undecided voters break to the challenger by better than 2-1, ESPECIALLY WHEN THE INCUMBENT IS UNPOPULAR, AS BUSH WAS IN 2004 (48.5% rating) and in 2006 (33% rating). That’s why I conservatively assume that Kerry in 2004 and the Democrats in 2006 would win the undecided vote by 60-40% in my election models. And that is why my 2004 PROJECTION (51.8 Kerry-48.2
Bush) and 2006 projection (57D-43R) WERE BOTH RIGHT ON THE MONEY. They were based on 18 final national pre-election polls and 116 pre-election Generic polls in 2006.


FEBBLE:
That incumbents below 50% always lose.

TIA
I never said ALWAYS. But Bush had 48.5% ratings in 2004. He stole the election. Carter (1980), Bush (1992), Ford (1976) all had ratings below 50%. And they all lost.

FEBBLE:
That Likely Voter models are less reliable than Registered Models.

TIA
I said they were less reliable in 2004, when new Democratic registrations were massive and young, single cell-phone users were unlikely to be contacted. These were NOT likely voters. Facts, Febble. Facts.

Fianally, you have claim elsewhere in this thread that my analysis is not supported in the "reality-based" community. Well, what about Freeman, Mark C. Miller, RFK,
Baiman, Dopp, Jonathan Simon, Bruce O'Dell, Conyers, Fitrakis, Richard Hayes Phillips, Palast, Michael Keefer,etc? They aint exactly chopped liver.

On The Other Hand, they aren't members of your "reality-based" community which includes E-M, Mystery Pollster, Rick Brady, Farhad Manjoo, Diebold, ES&S.

Febble (1000+ posts) Mon Nov-20-06 07:48 PM
Response to Reply #55
58. Last post

Edited on Mon Nov-20-06 07:58 PM by Febble
FEBBLE:
That sampling error is the only error in a poll.

TIA:
I never said that. I said that scientific polling minimizes the error and that’s why pollsters always quote the MoE based on sample-size.

Well, that isn't true. It isn't why pollsters always quote the MoE based on sample-size.

Here is Mitofsky on the subject:

REPORTING SAMPLING ERROR
I want to say a few words about reporting sampling error. A number of people who have spoken here have talked of not reporting sampling error because it was confusing all those dear mindless souls who listen to our results. They were concerned we would make people think that sampling error was the only error in the survey.

http://www.nyaapor.org/WMitofskySpeech.htm

You cannot "scientifically" eliminate bias from a survey. Inferring that from the fact that a pollster quotes the MoE based on sample size put TIA in the category of those "dear mindless souls" who would be "confused" into thinking such a thing.


FEBBLE:
That when there IS bias in the poll it would favor Republicans, not Democrats

TIA:
A few FACTS.
1) It’s a well known fact that approximately 3% of the votes are uncounted due to lost, spoiled and provisional ballots. The vast majority of these votes are in minority districts (50% in black districts). Since the minority black vote is 90% democratic, and the Hispanic vote at least 60%, a fair estimate is that 75% of the uncounted votes (2.25%) are Democrat and the other 0.75% Republican. That’s a 1.5% bias in favor of the Republicans.

2) There is a proven reluctance of low-income Democratic voters (under $50K) and high-income Republicans ($100k+) to participate in exit polls. But since low income voters outnumber high-income voters by almost 3-1, there's another component of Republican bias.


Spoiled votes would not cause bias in the sample. They would simply cause a discrepancy between the sample and the count. However, the effect is likely to be small because the spoiled votes tend to be concentrated in strongly Democratic precincts which are not highly represented in the precinct sample. It will be an effect however, although much smaller than the effect on the actual results. I share with TIA his indignation at this systematic disenfranchisement of largely Democratic voters. As for his second point - TIA read it in a book somewhere. It is not something that can, or should, be generalised to exit polls. There are many sources of evidence, including direct experimental evidence that in exit polls, Democratic voters tend to be over-sampled.

3) Exit poll response was high in Bush states and low in Kerry states.
Here's GRAPHIC MATHEMATICAL PROOF USING LINEAR REGRESSION ANALYSIS:

Overall response rates are not a proxy for response bias. Response bias occurs when the response rates for one set of voters differ from the response rate for the other set of voters, whether the two rates are 15% and 20% or 60& and 80%. Moreover, selection bias will not show up in response rates - and may even be associated with higher response rates. If more willing voters are being selected, completion rate will go up.


FEBBLE:
That people recall their previous votes correctly (there is excellent evidence that they do not, and that they tend to misreport having voted for the previous winner)

TIA:
If you consider that the FINAL National Exit poll is ALWAYS MATCHED TO THE RECORDED VOTE, WHICH MEANS THAT THE WEIGHTS ARE FUDGED TO MATCH THE VOTE. I CALL IT FUDGING; YOU CALL IT VOTER MISREPRESENTATION. Witness the 2004 Final NEP, which Bush won by 51-48%: 43% of 2004 voters were Bush voters and 37% were Gore voters? This was a FUDGE necessary to match the recorded vote. In the 12:22am poll, which Kerry won by 51-48%, the mix was 41 Bush/39 Gore. Mathematically, the Bush MAXIMUM Bush 2000 representation weighting was 39.8%, which is the ratio of Bush 2000 voters still alive n 2004 by the 122.3mm who voted.

Do you want MORE of this evidence? Look at the 2006 NEP. The 7pm poll had the 2004 weighting as 45Bush/46 Gore. The FINAL had it 49 Bush/43 Gore/8 Other. Where did the 8% for OTHER voters come from? Third-party 2004 voters comprised 1% of the vote. Where did the excess 7% come from? Kerry. Here’s why:

Are we to believe Bush voters outnumbered Kerry voters by 6%? Even if you believe the 2004 FINAL, which we have proven bogus, the Bush margin was 3%. In reality it should have been 50 Kerry-47 Bush, after deducting 1% for voter mortality over
the 2-year period. Now 50% = 43% + 7%. There is your Kerry vote.


I really can't be bothered to explain this to TIA again. He needs to read Mark Lindeman's paper:

http://inside.bard.edu/~lindeman/too-many.pdf

That undecideds always break for the challenger.

TIA:
I never said ALWAYS. I said the vast majority of the time. The evidence is a study of 155 incumbent elections that you yourself quoted: in 82%, the challenger won the undecided vote; the incumbent, just 12%; neither, 6%. And of course, pollsters Zogby, Harriss and others have always claimed that undecided voters break to the challenger by better than 2-1, ESPECIALLY WHEN THE INCUMBENT IS UNPOPULAR, AS BUSH WAS IN 2004 (48.5% rating) and in 2006 (33% rating). That’s why I conservatively assume that Kerry in 2004 and the Democrats in 2006 would win the undecided vote by 60-40% in my election models. And that is why my 2004 PROJECTION (51.8 Kerry-48.2
Bush) and 2006 projection (57D-43R) WERE BOTH RIGHT ON THE MONEY. They were based on 18 final national pre-election polls and 116 pre-election Generic polls in 2006.


Well, did he weight his probability by his estimate of the probility of his assumption being true?

FEBBLE:
That incumbents below 50% always lose.

TIA
I never said ALWAYS. But Bush had 48.5% ratings in 2004. He stole the election. Carter (1980), Bush (1992), Ford (1976) all had ratings below 50%. And they all lost.


FEBBLE:
That Likely Voter models are less reliable than Registered Models.

TIA
I said they were less reliable in 2004, when new Democratic registrations were massive and young, single cell-phone users were unlikely to be contacted. These were NOT likely voters.Facts, Febble. Facts.


And TIA didn't investigate the fact that this was compensated for by up-weighting the age demographic.


Finally, you have claim elsewhere in this thread that my analysis is not supported in the "reality-based" community. Well, what about Freeman, Mark C. Miller, RFK,
Baiman, Dopp, Jonathan Simon, Bruce O'Dell, Conyers, Fitrakis, Richard Hayes Phillips, Palast, Michael Keefer, etc? They aint exactly chopped liver.

They ain't exactly unanimous either.

OK, I'm calling this off. I can't converse with a poster who isn't here. If TIA wants me , he knows where to find me.

Peace.

Lizzie

Febble (1000+ posts) Sat Nov-18-06 04:31 PM
Response to Reply #36
37. Sure

And like you guys with your DINOs you do what you can to get them replaced. It was one of the reasons I was desperate for a Kerry win - it would have been a great boost to the anti-Bush faction within the government. I hope Gordon Brown will take over within a few months.

Regarding TIA's work: this piece of work seems largely to do with his own pre-election predictions which were very much more optimistic than those of most pundits, and was based on generous assumptions.

Try:
http://www.centerforpolitics.org/crystalball /
and check out the plots here:
http://www.pollster.com/polls /

With polls, your answer depends on your assumptions. As you cannot be sure your assumptions are correct, you should also estimate the probability that you are correct (a judgement call). TIA does not do this. He also assumes that the only error in the exit poll is sampling error, despite good evidence that exit polls in the US tend to show a pro-Democratic bias (he does not believe this is the case).

See
http://www.dailykos.com/storyonly/2006/11/4/135126/905
for more on this, by me.

I have no problem with TIA's calculations, but when you compute probabilities you make certain assumptions. If these are violated, or not justified by the nature of the data, then your answer will be wrong. I think TIA's answer is wrong, for those reasons.

I'm not sure what you mean by reading between pronouns, but I have now explained my use of the second person.


mom cat (1000+ posts) Sat Nov-18-06 05:32 PM
Response to Reply #37
39. Response from TIA

Febble, you obviously ignored the model analysis. In fact, your lack of focus betrays your fixed agenda once again: to debunk any and all analysis which point to election fraud is based on pre-election and exit polling data.

You are wrong. I do take Democratic exit poll bias into account:
1) Of the 3% of uncounted votes in every election, the vast majority are Democratic. That would skew the exit polls to the Democrats. I'm surprised you still don't get it: UNCOUNTED DEMOCRATIC VOTES WOULD EXCEED UNCOUNTED REPUBLICAN VOTES BY 1.5% (2.25%-0.75%R). OF COURSE, THE DEMOCRATIC EXIT POLL RESPONDENT HAS NO WAY OF KNOWING THAT HIS/HER VOTE WAS NOT COUNTED.

2) I certainly do take probabilities into account in my projections via Monte Carlo simulation of 1000 trial elections using each of the 61 pre-election polling data and a 3% MoE.

3) The Undecided Voter Allocation assumption that a majority (60% is conservative) break for the challenger is based on the historical study of 155 incumbent elections which you are well aware of and in fact mentioned in one of your posts. All 61 incumbents in this analysis were Republicans.

4) What do you have to say about the 116 pre-election generic polls, all won by the democrats, in which the trend line pointed to a 14% democratic margin, after applying a 60% UVA? Why don't YOU calculate the probability that the final results would be off by at least 3X the MoE?

5) And finally, how do you explain the 49%Bush-43%Kerry Voted in 2004 weightings? Like you explained the 43% Bush/37% weights in the Final 2004 NEP?How do you explain the remaining 8% when onlt 1% voted for third party candidates in 204? You can't logically, but you will surely try. Are you going to say that voters were more 6% likely to say they voted for Bush with his 35% rating?

Febble, the reality based FACTS are NOT on your side.



Febble (1000+ posts) Sat Nov-18-06 05:59 PM
Response to Reply #39
40. Response to TIA

1. I do not have a fixed agenda. This false.

2. I am quite aware that uncounted Democratic votes may contribute to a discrepancy in the exit polls. Indeed, in my work for Mitofsky I specifically looked for evidence of this, and found some that, while not conclusive, is suggestive that it may have played a role. The measure of precinct level discrepancy I used, developed together with Mark Lindeman, meant that it was more sensitive than the traditional measures to bias in extreme precincts, and this may have been why my analysis was able to detect something. The finding was that in urban precincts, particularly largely black urban precincts, those precincts in which older technology was used (levers and punchcards) the discrepancy was greater then where digital technology was used. It was a small effect, and may have been confounded by other factors, but it was of interest. However, because uncounted Democratic votes tend to be concentrated in extreme Democratic precincts, it is unlikely that even large numbers of residual votes would have a large impact on the exit poll, simply because extreme Democratic precincts are not heavily represented in the precinct sample. I agree with TIA, and with Greg Palast (though I would still like details of where he gets his numbers) that uncounted Democratic votes, together with voter suppression, cost the Democrats many votes in each election, and, indeed, cost Gore the presidency. But the exit polls are unlikely to reflect this loss except at precinct level, and discrepancies between overall exit poll estimates and counted results are unlikely to index, or even reflect, the magnitude of this problem.

3. What I said is that TIA does not factor in the probability that his own assumptions are in error.

4. As many others have argued, translating a generic ballot into seats is difficult, given the nature of congressional district boundaries, and, indeed, the number of seats up for election. Nonetheless several pundits attempted it. None that I was following came up with TIA's optimistic projections. Larry Sabato pretty well nailed it.

5. As TIA knows, retrospective inflation of the winners margin is a routine phenomenon in exit polls, regardless of the popularity of the incumbent. It even happened with Nixon, who was not even president by then. These facts are not on TIA's side.


mom cat (1000+ posts) Sun Nov-19-06 03:33 AM
Response to Reply #40
43. TIA response to FEBBLE and other interested DUers.

DUers may not be familiar with these facts.
Given:
Bush 2000 vote: B= 50.45mm
2004 total vote: N= 122.3mm
Annual U.S. mortality rate: R= 0.87%
Final NEP: W= Bush 2000 voters/N = 43%

Calculate:
1 The number of Bush 2000 voters (D) who died prior to 2004: D = 4*R*B
2 The maximum number (X) of Bush 2000 voters who could vote in 2004: X = B–D
3 The maximum weighting (W) of Bush 2000 voters who voted in 2004: W = X/N

If the calculated W does not equal the Final NEP W of 43%, which one is correct?

What does this tell us about the "false recall" theory, that a significant percentage of exit poll respondents falsely report who they voted for in the prior election?

What does this tell us about the Reluctant Bush Responder (rBr) theory, that for various reasons Democrats are more likely to be exit-polled than Republicans?


Febble (1000+ posts) Sun Nov-19-06 08:10 AM
Response to Reply #43
44. Well, TIA knows the answer to this one
or if he doesn't, it's not for want of being told.

If the calculated W doesn't equal the Final NEP W, then the probability is that, as usual, people misreported their previous vote in favor of the winner.

This "retrospective inflation" of the winner's margin is evident, as Mark Lindeman showed here:

http://inside.bard.edu/~lindeman/too-many.pdf

in every single presidential exit poll since 1976, and applies whether the previous winner is Republican, Democrat, running, losing, or even still in office (Nixon). The fact that Bush's retrospective margin was not much inflated in the unadjusted poll, is if anything, evidence that adjustment was required - that Bush 2004 voters had been undersampled. The adjusted margin is more inline with the kind of retrospective margin inflation expected. TIA makes exactly the same error that O'Dell and Simon just made in their recent "Landslide Denied" paper.

So that is the answer to TIA's question. It tells us that Kerry voters were probably sampled at a higher rate than Republicans (as is evident from other analyses) and that people misreported their previous vote in favor of the winner in the same kinds of proportions as would be expected, given the pattern we observe in the historical data.

TIA needs get his head out of Excel and into Adobe Acrobat.


mom cat (1000+ posts) Sun Nov-19-06 09:13 PM
Response to Reply #44
47. Feeble. Once again you avoid the facts.

Febble, once again you stubbornly avoid the FACTS and the IMPLICATIONS of the FACTS.

First of all, you didn’t do the math. Because if you did, you would know that at MAXIMUM, only 48.7mm Bush 2000 voters could have voted in 2004. And therefore the MAXIMUM Bush weighting was 39.8% = 48.7/122.3%. The 43% weighting was IMPOSSIBLE. This is a MATHEMATICAL FACT.

Therefore, if you were willing to accept this MATHEMATICAL FACT, you would have to agree that the Final NEP 43% Bush weighting was a FICTIONAL ARTIFICE AND NOT A SAMPLED RESULT. IT WAS AN ARBITRARY FUDGE WHICH WAS REQUIRED TO FORCE THE FINAL NEP TO MATCH THE BUSH RECORDED VOTE.

Even you have agreed that the Final NEP is ALWAYS matched to the recorded vote.

Why, for heaven's sake, is it not yet clear to YOU that matching to the recorded vote count ONLY makes sense IF the vote count is ACCURATE and there is ZERO FRAUD? This should be clear to everyone, even a third grader who cheats on his arithmetic test.

THEREFORE, THE FINAL NEP IS A FRAUD.

When will you accept the FACT that the RECORDED 2004 VOTE was BOGUS and that BushCo used massive FRAUD to STEAL the election? In fact, BushCo has successfully stolen EVERY election since 2000 -except for 2006. They were stopped in 2006 only because of the Democratic TSUNAMI.

The reality-based community must therefore conclude that your “retrospective” exit poll argument HOLDS NO WATER AND IS JUST ANOTHER RUSE TO DEFLECT FROM THE FINAL NEP'S EGREGIOUS MATCHING TO FRAUDULENT, MISCOUNTED VOTES.

If you were a true analytical investigator, you would not employ TWISTED LOGIC TO DEFEND THE FINAL NEP AND CONSTANTLY CRITICIZE THE ACCURACY OF PRE-ELECTION AND PRELIMINARY EXIT POLLS WHICH CLEARLY POINT TO FRAUD.

If you were a true analyst, you would stick to the mathematical facts.


OnTheOtherHand (1000+ posts) Sun Nov-19-06 09:28 PM
Response to Reply #47
48. why this is wrong

Short version: as has been explained many times, there is no reason to expect exit poll respondents to report their past votes accurately. Therefore, while it may be true that no more than 39.8% of 2004 voters could have voted for Bush in 2000, there is no clear limit to what percentage of 2004 voters could have claimed to vote for Bush in 2000. The apparent overstatement is in line with overstatements of previous winners' vote shares in other exit polls.

It must be frustrating to believe that the entire reality-based community should have looked at the exit poll tabulations in November 2004 and realized that Kerry won the election. All those people blithely writing as if Bush won -- are they innumerate? careless? craven?

Maybe they just don't take polls as literally as, apparently, you do. Skepticism about polls could be reality-based, don't you suppose?


Febble (1000+ posts) Mon Nov-20-06 02:35 AM
Response to Reply #47
51. mom cat:

Is this you talking, or am I talking to TIA? Either way, please get my name to right.

Only the first of those is anything close to a mathematicial FACT. They are mathematical inferences, and assertions based on a stubborn refusal to understand the nature of the data.

I have attempted repeatedly to explain to TIA why his inferences are erroneous, and he is clearly as frustrated as I am with him that I don't see it his way.

However, if these arguments are yours, I am happy to discuss them with you.


mom cat (1000+ posts) Mon Nov-20-06 10:08 AM
Response to Reply #51
53. I am sorry about the misspelling of your name. It was a mistake,
not intentional.


Febble (1000+ posts) Mon Nov-20-06 11:18 AM
Response to Reply #53
54. No problem

It's TIA's assumptions I have a greater problem with!


mom cat (1000+ posts) Mon Nov-20-06 07:06 PM
Response to Reply #51
56. I thought for sure that you would recognize that the post was from TIA.



Febble (1000+ posts) Mon Nov-20-06 07:25 PM
Response to Reply #56
57. Well, I thought it probably was!

He has an inimitable style!