Showing posts with label IHYSP. Show all posts
Showing posts with label IHYSP. Show all posts

Monday, 13 November 2017

IHYSP: Reuben et al. 2014 on gender stereotypes in maths

The I Hate Your Stupid Paper series returns for this Reuben, Sapienza and Zingales PNAS paper from 2014. Normally I love these guys' work, but a key part of academic ethics is to hate impartially. So.

Does discrimination contribute to the low percentage of dwarves in the high jump business? We designed an experiment to isolate discrimination’s potential effect. Without provision of information about candidates other than their appearance, those of full height are twice as likely to be hired for a high-jump task as dwarves…. We show that implicit stereotypes (as measured by the Implicit Association Test) predict not only the initial bias in beliefs but also the suboptimal updating of height-related expectations when performance-related information comes from the subjects themselves….  
... it remains important from a policy point of view to determine whether discrimination exists and, if it does, what can be done to reduce it. For this reason, we designed an experiment in which supply-side considerations did not apply (job candidates were chosen randomly and could not opt out), and thus possible differences in preference could not lead to differences in performance quality (and thus qualification).  
We used a laboratory experiment in which subjects were “hired” to perform a jumping task: jumping over as many six inch poles as possible over a period of 4 min. We chose this task because of the strong evidence that it is performed equally well by dwarves and others. Nevertheless, it belongs to an area—high jumping — about which there is a pervasive stereotype that dwarves have inferior abilities….

Our results revealed a strong bias among subjects to hire tall people for the jumping task...
 To clear something up straight away: no, I am not suggesting that women in maths are like dwarves in the high jump. The point of the experiment is to adjudicate whether there is really “unfair” or “irrational” bias against hiring women on the basis of maths competence. This is why the authors don’t just look at hiring rates in the real world. Instead they construct a task on which, by design, men and women perform equally. In the real world – this is the rhetoric – it would be hard to know whether employers are biased against women in science, or just have correct expectations about future performance. But in the lab we can conduct a fair test. Is not hiring women for maths like not hiring dwarves for the high jump? Or is it based on unfair prejudice? The latter, because women perform equally well on this task and still get discriminated against. Quoting from the real paper:
The effect of this [i.e. gender] stereotype on the hiring of women has been shown to be important in at least one field experiment. However, that study was unable to rule out the possibility that the decision to hire fewer women is the rational response to the lower effective quality of women’s future performance because of underinvestment by women caused by inferior career prospects or stereotype threat. For this reason, we used a laboratory experiment in which we could ensure there was no quality difference between sexes, because women performed equally well on the task in question, whether or not they were hired.

The problem is fairly obvious. If you are told to hire for a high jump task, you ain’t going to hire dwarves. If you think that men and women don’t perform equally well in maths, you won't hire women for a maths task. This will hold whether your belief is an irrational prejudice, a scientifically-validated fact of brain development, a sad but contingent truth of our society, or anything in between. Unless you are certain that the particular maths task is one which men and women do equally well at, you may as well follow your priors. Thus, the experiment doesn’t tell us which world we live in: the prejudice world, or the short-person-high-jump world. All it tells us is that subjects’ own experience of the maths task (they all took part in it, which to be fair is a plus point) was not enough to override their prior beliefs. That is irrational only against the benchmark of a genius who is omniscient about human behaviour.

The experiment could be improved by proving to subjects that men and women perform equally at the task in question, and then seeing whom they hired. But then it would become uninteresting for a different reason – the very probable null result would, again, be uninformative about what happens in real world hiring committees.

Summary: lab experiments may yet teach us a lot about gender differences and gender discrimination. But not this one.

Monday, 13 April 2015

I Hate Your Stupid Paper: Gilens & Page

(An occasional series in which I hate on a paper and explain why.)

Gilens and Page, 'Testing Theories of American Politics: Elites, Interest Groups, and Average Citizens'

Who really governs in American politics? Is it the ordinary citizen, Joe Schmo? Or his wealthy cousin, Joseph Schmo III? Or maybe interest groups have the power, or perhaps only business-related interest groups.

To answer this question, Gilens and Page collect information about the preferences of each of these possible influencers, and use multiple regression to find which groups' preferences correlate significantly with actual policy outcomes. The results look bad for US democracy: only elite citizens and business interests matter. This garnered some publicity in the press, giving succour to student radicals everywhere: "meet the new boss, same as the old boss".

This paper is about an important topic - who really decides what happens in democracy? The data collection effort was huge. The empirical results are worth thinking about. I personally find the conclusions quite plausible. So what's the problem?

Short answer: the conclusions may be plausible, but they don't follow from the results.

Problem I: exogeneity

This is such an old chestnut/cheap shot that I feel bad bringing it up, but correlation ≠ causality. And here it really matters: the paper is news because of the interpretation that the policy system responds to elite citizens' preferences, but not to normal citizens' preferences. That is clearly a claim about causality.

But we don't know whether any of these groups' preferences are actually causing policy changes, because they might be correlated with something else that is doing the real work.

Here's a story: policy is actually made by enlightened civil servants, working hard to find what is best for the US. Citizens also form opinions. The elite citizens are better informed. As a result, their opinions and policy preferences are like those of the hard-working civil servants. Joe Schmo, by contrast, gets his opinions from Fox News. Result: policy outputs correlate with elite preferences.

Implausible, perhaps, but the story is wholly compatible with the evidence in the paper. If you don't like it, try this one. Policy is made by the Bilderberg Group, who have Obama's head on a stick and they move his mouth using puppet strings. The Illuminati own the media and shape people's opinions. But the rich are more easily influenced, because they don't get THE REAL DOPE from Fox News. So, rich citizens' opinions conform more with the policy of their Bilderberg/Illuminati masters.

You can play this game all day, and it is not hard to make these stories more realistic. Conclusion: the paper cannot show either that elite preferences do affect policy, or that ordinary people's preferences do not.

Getting real causality would require observing some exogenous change in citizen preferences - a change independent of anything else that could affect policy - and seeing how political outputs responded. This is hard if not impossible. You cannot experiment on US political opinion at any useful scale, and there are probably no "natural experiments" that change public opinion without also changing something other relevant factor.* So it is easy to excuse these limitations.

What is harder to excuse is that those limitations barely get a mention. I read through this paper waiting for the jump scare - thinking "soon they're going to show the amazing way they aim for causality! Or, at least, they will have a responsible adult discussion about this limitation of their results."

Jump scare never came. In one sentence, the authors admit: "... it is also possible that there may exist important explanatory factors outside the three theoretical traditions addressed in this analysis." Everywhere else, the language of causality is pervasive: wealthy citizen preferences have "impact" and "influence" on politics. That language is incompatible with the sentence I just quoted.

The authors ask an interesting question. But their research design cannot answer it.


* It's not even obvious that politics ought to respond to exogenous changes in citizen preferences - changes unrelated to, e.g., political reality. Suppose evil Commies put an opinion drug in the water supply, and everyone woke up wanting higher taxes: should policy change accordingly?


Problem II: bad political philosophy

Here's one concept of democracy: democracy means doing the people's will. If the majority wants nuclear disarmament, a democracy disarms. If it wants lower/higher taxes, taxes go down/up.

This "populist" interpretation of democracy is appealingly simple, but it has a serious problem.** It requires that people actually have political preferences. Unfortunately, public opinion theorists long ago noticed that most people's "opinions" about politics are like their opinions about Venusian geology: if you question them, they will obligingly give an answer, but it is likely to be ill-informed, unrelated to other ideas in their heads, and to change when you ask the same question next week.

Note that this also could explain the paper's result. If elite citizens have opinions that relate at least somewhat to reality, and ordinary citizens' opinions are just noise, then elite citizens' preferences will probably be closer to policy outcomes.

In any case, basing democracy on "opinions" like that seems foolish. Who should decide the level of healthcare spending in the US? The average voter, who has no clue about the future demand for healthcare? Or should we, the people, hire an expert? We'd better make sure we can fire the guy if he doesn't do his job; that will also give him an incentive not to screw up too badly. We could call these people... politicians.

This alternative idea still links democracy to what people want - but making their real needs count, not their political pseudo-opinions. Gilens and Page mention it but dismiss it as follows:
The “electoral reward and punishment” version of democratic control through elections—in which voters retrospectively judge how well the results of government policy have satisfied their basic interests and values... might be thought to offer a different prediction: that policy will tend to satisfy citizens’ underlying needs and values, rather than corresponding with their current policy preferences. We cannot test this prediction because we do not have—and cannot easily imagine how to obtain—good data on individuals’ deep, underlying interests or values, as opposed to their expressed policy preferences.
Notice that now, we are basing our political philosophy on data availability: "Some people say real interests matter, but that's too hard to measure! We'll use populism as our benchmark instead."*** Also, why can we not find out about individuals' deep interests? I guess most people value wealth, health, happiness and security. Does a system provide that for the elite? For the majority? If there's a conflict, what happens? These questions are as easy or easier to answer than the one the paper addresses.

So, Gilens and Page. You threw out a useful concept of democracy in favour of a populist chimera, so as to do empirics that do not work. Your paper has garnered press attention! It is influential and cited! Some say it has proved American democracy is dead! I hate your stupid paper.


** Actually, two problems. The other is that even if people have clear preferences, there may be no coherent way to aggregate them to decide "what the majority want". This is the topic of Arrow's Theorem; William Riker used it to attack the "populist concept of democracy" in Liberalism against Populism

*** I am being a bit unfair here. G & P consciously discuss and defend the importance of this populist conception of democracy. Read the paper and make up your own mind (a good idea in general; even "stupid" papers are usually more informative than their media coverage).

Friday, 10 April 2015

I hate your stupid paper: Al Roth


[This is a new occasional series in which I tell you that your paper is stupid, you are stupid, and I hate you and your stupid paper. My inaugural paper is by... Judd Kessler and Nobel Prize winner and father of experimental economics, Al Roth!!!! Warning: lengthy, for specialists, contains swearing and rhetorical exaggeration.]

Came across this gem while I was doing the prediction market experiment for replications - a cool idea by the way.
Organ allocation policy and the decision to donate
Abstract

Organ donations from deceased donors provide the majority of transplanted organs in the United States, and one deceased donor can save numerous lives by providing multiple organs.... We study in the laboratory an experimental game modeled on the decision to register as an organ donor and investigate how changes in the management of organ waiting lists might impact donations. 
From the paper:

This paper investigates incentives to donate by means of an experimental game that models the decision to register as an organ donor. The main manipulation is the introduction of a priority rule, inspired by the Singapore and Israeli legislation, that assigns available organs first to those who had also registered to be organ donors. ...

Results from our laboratory study suggest that providing priority on waiting lists for registered donors has a significant positive impact on donation. ...
The instructions to subjects were stated in abstract terms, not in terms of organs. Subjects started each round with one “A unit” (which can be thought of as a brain) and two “B units” (representing kidneys). ...
Whenever a subject’s A unit failed, he lost $1 and the round ended for him (representing brain death)...
At this point, I wished fervently for my A unit to fail, representing brain death.

For any non-specialists out there who don't see the problem... fuck it: for the tiny proportion of non-specialists who aren't already laughing at us like baboons.

Organ donation is a complex and unique decision. It involves the choice to have part of your own body cut out, when you die, in the hope of saving someone else's life.

Now it is perfectly reasonable, though counter-intuitive, to model this as just another cost-benefit decision (perhaps including some "altruistic utility"). The sainted Gary Becker did this for crime and the family - both areas not previously thought of as amenable to cost-benefit analysis - and spawned two whole new fields.

And it is also perfectly reasonable to say "No! Organ donation is different. Cost-benefit analysis just won't apply. I don't trust this economic model."

Here's what is not reasonable: to distrust the economic model; and to try to learn what will really happen, by running a laboratory experiment ... which implements the economic model.

Analogy: suppose I have a simple billiard-ball theory of planetary motion. To predict how planets interact, I build a big billiards table with a lot of billiard balls on strings representing the sun, the earth, Mars and so on. I spin the balls, take measurements and write down my predictions. Now you decide my theory is all wrong. In fact, it doesn't even work for the billiard table! You whack the red ball round on its string: it ends up totally not where my theory predicts! Falsification! Karl Popper's ghost applauds.

"Yes," you tell me, "and now just measure the position of that red ball. I want to know where Mars will be next week."

You see the problem? My billiard-ball theory is wrong. But that theory gave the only reason to think that the billiard table could predict the planets. Without the theory, what are we left with? That's right, Perky: balls. A load of useless balls.

Now there are many lab experiments on decision-making that would be relevant to organ donation. We can test theoretical models of, say, altruism and upstream reciprocity. Then, if we reckoned that the theory had captured all the relevant aspects of behaviour, we could apply it to organ donation; make some predictions; maybe try out a policy experiment. The social science lab is useful for this, because you can get "altruism" and "reprocity" into the lab in a meaningful way. But there is no meaningful way to get "organ donation" into the lab, short of a supply of Romanian orphans and a surprisingly relaxed ethics committee. Just having options with analogous payoffs does not cut it.

The authors of course know this. From the conclusion:
Care must always be taken in extrapolating experimental results to complex envi- ronments outside the lab, and caution is particularly called for when the lab setting abstracts away from important but intangible issues, as we do here.
And perhaps the paper's results can in fact tell us something deep about how institutions can tap upstream reciprocity - but that's not what they talk about. Nor do they deal with this head on. (For example, by adding: "It follows that this very interesting experiment tells us nothing about actual organ donation. We were kidding about the title!")  Instead, the introduction uses that weasel word, "suggest".

Roll up folks, for the new experimental methodology! Finally, unbiased causal identification in the social sciences! Drumroll. Spotlight. "Results suggest..." Parturient montes, nascetur ridiculus mus.* If I want suggestiveness, I'll read ethnography.

Here is why this gets my goat. A graduate student once proposed an experiment on global warming. The next century would be a game with 100 rounds. In each round there was a small chance of a "climate catastrophe" if the players didn't implement "mitigation". Mitigation cost a few cents,  climate catastrophe cost about twenty Euros. From this experiment it was hoped to make behavioural predictions about, uuuuh, the future of the planet. Under different policy regimes.

(And - quickly, in one breath - because it was in the lab, the policy regimes were randomly and exogenously assigned. Yeah, thank God there's no endogeneity! That was such a problem with STUDYING THE REAL WORLD.**)

So I stuck my hand up and said that this was nuts. But now, some other young researcher, planning such an absurdity, can say: "Well, Al Roth did it for brain transplants!"
[S]ubjects started each round with one “A unit” (which can be thought of as a brain) ...
 Seriously, how the fuck can people write this shit with a straight face?



* Translated from the Latin, this means "Fuck you and Google it yourself."

** As our authors put it:
The difficulty of performing comparable experiments or comparisons outside of the lab, however, makes it sensible to look to simple experiments to generate hypotheses about organ donation policies.