Showing posts with label Michelle Amazeen. Show all posts
Showing posts with label Michelle Amazeen. Show all posts

Thursday, June 14, 2018

Different Strokes for Different Quotes: What does "voted for tax cuts" really mean?

"What I find is it's hard for me to take critics seriously when they never say we do anything right. Sometimes we can do things right, and you'll never see it on that site."

-PolitiFact Editor Angie Drobnic Holan



Sometimes PolitiFact can do things right.

PolitiFact New York did something right recently that deserves mention because it's the correct way to journomalist:





PolitiFact added the Trump camp "did not get back to us with information supporting his claim, so we can't say for sure what he was talking about in his endorsement."

PolitiFact noted that Trump tweeted about the Tax Cuts and Jobs Act "four other times in May" but acknowledged Trump did not reference that law in the tweet it fact checked.

In our view this is the correct approach.

We think a persuasive argument could be made that Trump inaccurately implied Donovan was a Tax Cut and Jobs Act supporter, but that argument belongs on the editorial page, not in a fact check. PolitiFact examined the claim Trump made without inventing assumptions about what he meant or what he was implying. In this case PolitiFact stuck to the facts.

Notwithstanding our longtime opposition to rating facts on a sliding scale, we think PolitiFact did this one right and we're happy to point it out.


The Other Guy

Readers may wonder "How could a fact checker screw this one up?" Donovan had a documented history of voting for tax cuts, and Trump's claim was not only unambiguous but also easy to check.

How could a serious fact checker get this wrong?







When the Washington Post's unabashed Trump basher/unbiased truthsayer tweeted that "fact checkers sometimes disagree" we were curious. PolitiFact rated Trump's tweet as accurate, while Kessler deemed the exact same tweet false. How can that be?

As it turns out the two fact checkers aren't disagreeing at all.

PolitiFact correctly identified the claim Trump made and ruled based on his actual words. Kessler invented a claim and then gave Trump a false rating for his own fantasy. The fact checkers aren't disagreeing because they're not checking the same claim.

Kessler says Trump's claim that Donovan "voted for tax cuts" is false because "Donovan voted against Trump's tax cut three times." For those of you that aren't experts in journalism or logic, voting against the Tax Cut and Jobs Act does not negate the fact that Donovan has previously voted for other tax cuts.

As far as we can tell, Kessler offered no justification for calling Trump's claim false other than Donovan's opposition to the 2017 tax bill.

Kessler's reasoning here is flatly wrong. And if one wanted to treat Kessler with the same painful pedantry as he applies to Trump in his chart, one could note there's no such thing as "Trump's tax cuts" because only Congress can pass tax bills.

Petty word games aside, this "disagreement" among fact checkers affirms that our fact-divining betters are neither scientific agents of truth nor objective determiners of evidence. When a fact checker can substitute a person's actual words for their own interpretation of what that person meant it counts as commentary, not an adjudication of facts.

Kudos to PolitiFact New York for taking the correct approach. Sometimes PolitiFact can do things right.


Friday, May 19, 2017

What "Checking How Fact Checkers Check" says about PolitiFact

A study by doctoral student Chloe Lim (Political Science) of Stanford University gained some attention this week, inspiring some unflattering headlines like this one from Vocativ: "Great, Even Fact Checkers Can’t Agree On What Is True."

Katie Eddy and Natasha Elsner explain inter-rater reliability

Lim's research approach somewhat resembled research by Michelle A. Amazeen of Rider University. Amazeen and Lim both used tests of coding consistency to assess the accuracy of fact checkers, but the two reached roughly opposite conclusions. Amazeen concluded that consistent results helped strengthen the inference that fact-checkers fact-check accurately. Lim concluded that inconsistent fact-checker ratings may undermine the public impact of fact-checking.

Key differences in the research procedure help explain why Amazeen and Lim reached differing conclusions.

Data Classification

Lim used two different methods for classifying data from PolitiFact and the Washington Post Fact Checker. She converted PolitiFact ratings to a five-point scale corresponding to the Washington Post Fact Checker's "Pinocchio" ratings, and she divided ratings into "True" and "False" groups using the line between "Mostly False" and "Half True" as the barrier between true and false statements.

Amazeen opted for a different approach. Amazeen did not try to reconcile the two different rating systems at PolitiFact and the Fact Checker, electing to use a binary system that counted every statement rated other than "True" or "Geppetto check mark" as false.

Amazeen's method essentially guaranteed high inter-rater reliability, because "True" judgments from the fact checkers are rare.  Imagine comparing movie reviewers who use a five-point scale but with their data divided up into great movies or not-great movies. A one-star rating of "Ishtar" by one reviewer would show agreement with a four-star rating of the same movie by another reviewer. Disagreements only occur when one reviewer gives five stars while the other one gives a lower rating.

Professor Joseph Uscinski's reply to Amazeen's research, published in Critical Review, put it succinctly:
Amazeen’s analysis sets the bar for agreement so low that it cannot be taken seriously.
Amazeen found high agreement among fact checkers because her method guaranteed that outcome.

Lim's methods provide for more varied and robust data sets, though Lim experienced the same problem Amazeen found in that two different fact-checking organizations only rarely check the same claims. Both researchers used relatively small data sets.

The meaning of Lim's study

In our view, Lim's study rushes to its conclusion that fact-checkers disagree without giving proper attention to the most obvious explanation for the disagreement she measures.

The rating systems the fact checkers use lend themselves to subjective evaluations. We should expect that condition to lead to inconsistent ratings. When I reviewed Amazeen's method at Zebra Fact Check, I criticized it for applying inter-coder reliability standards to a process much less rigorous than social science coding.

Klaus Krippendorff, creator of the K-alpha measure Amazeen used in her research, explained the importance of giving coders good instructions to follow:
The key to reliable content analyses is reproducible coding instructions. All phenomena afford multiple interpretations. Texts typically support alternative interpretations or readings. Content analysts, however, tend to be interested in only a few, not all. When several coders are employed in generating comparable data, especially large volumes and/or over some time, they need to focus their attention on what is to be studied. Coding instructions are intended to do just this. They must delineate the phenomena of interest and define the recording units to be described in analyzable terms, a common data language, the categories relevant to the research project, and their organization into a system of separate variables.
The rating systems of PolitiFact and the Washington Post Fact Checker are gimmicks, not coding instructions. The definitions mean next to nothing, and PolitiFact's creator, Bill Adair, has called PolitiFact's determination of Truth-O-Meter ratings "entirely subjective."

Lim's conclusion is right. The fact checkers are inconsistent. But Lim's use of coder reliability ratings is, in our view, a little like using a plumb line to measure whether a building has collapsed due to earthquake. The tool is too sophisticated for the job. The "Truth-O-Meter" and "Pinocchio" rating systems as described and used by the fact checkers do not qualify as adequate sets of coding instructions.

We've belabored the point about PolitiFact's rating system for years. It's a gimmick that tends to mislead people. And the fact-checking organizations that do not use a rating system avoid it for precisely that reason.

Lucas Graves' history of the modern fact-checking movement, "Deciding What's True: The Rise of Political Fact-Checking in American Journalism," (Page 41) offers an example of the dispute:
The tradeoffs of rating systems became a central theme of the 2014 Global Summit of fact-checkers. Reprising a debate from an earlier journalism conference, Bill Adair staged a "steel-cage death match" with the director of Full Fact, a London-based fact-checking outlet that abandoned its own five-point rating scheme (indicated by a magnifying lens) as lacking precision and rigor. Will Moy explained that Full Fact decided to forgo "higher attention" in favor of "long-term reputation," adding that "a dodgy rating system--and I'm afraid they are inherently dodgy--doesn't help us with that."
Coding instructions should provide coders with clear guidelines preventing most or all debate in deciding between two rating categories.

Lim's study in its present form does its best work in creating questions about fact checkers' use of rating systems.