23 Comments
User's avatar
Victor Kumar's avatar

I would upgrade my subscription if the paywalled version named names

Jennifer Rogers's avatar

I would pay good money for this!

Christopher J Ferguson, Ph.D.'s avatar

Great essay and thanks for sharing.

I remember PsychMAP and it's coconspirator Psych Methods. I didn't stay long in either for the exact reasons you mention (I didn't know about the sock puppets though).

My only quibble though is you say "the reformers have won"....I guess that's *sorta* true, but did they really? I'm less sure.

My casual observation is stuff like preregistration is still not the norm (and, of course, many papers don't follow their preregs). Even the "reformers" ignore stuff like crud/noise and the self-defensive arguments around small effect sizes that are likely false positives.

I think there's been progress, but I see it more like the rebels at the time of Bunker Hill rather than at Yorktown.

I saw someone put it recently (paraphrased and I forget who): "It's the same bad science, just with bigger samples". I sadly think that's mostly true.

In the end, do I have much more confidence in psychological science today than I did 20 years ago? I'd say not much. the main "change" psychology seems to have made due to the replication crisis, is simply to lower standards of evidence. Where (in my field) people used to brag about r = .3 (likely due to p-hacking), people simply started bragging about r = .03 with little change in message. Until we fix that, I don't think we've accomplished much.

Jared Peterson's avatar

I was also in both PsychMAP and Psych Methods. The two blur together in my mind though, and I also didnt know anything about the sock puppets. Mostly I just remember that Uri guy. He had...a strong personality (coincidently, I believe he was a personality psychologist).

I'm still unsure whether the issue is that our methods are not rigorous enough, or whether the problem is the premise that we are discovering hidden truths about human nature that are replicable. An experimental average, albeit statistically different from a control group, seems like a pretty low bar for saying you have discovered something about human nature. Perhaps the human brain does not consist of a series of effects, tendencies, and biases? Perhaps the idea of stable dispositions modulated by context is just entirely the wrong way to think about the brain/mind? If so, then simply increasing the rigor will not solve the problem

Michael Inzlicht's avatar

I see replicability as the lowest possible bar for a science. If you repeat it, will you get the same result? And I believe--and I think there are audits that back this up--that we've improved on this front. But, you are right that this is far from what is needed to make good infereneces about the mind and the world we live in. For that, we need to work on external validity, measurement, generalizability a lot more. My sense is that people are doing this, but perhaps not as quickly as we'd like.

Christopher J Ferguson, Ph.D.'s avatar

I think so too, but I also think there's a lot of reluctance/pushback too. There are a lot of "small is big" arguments around effect size that are non-empirical (for instance the idea of accumulating effects over time, often refuted by longitudinal studies yet repeated anyway), and I've watched as effect size recommendations have gotten ridiculously generous (Funder & Ozer, 2019; Orth et al., 2024) such that, in the case of Ozer, we're literally interpreting r = .03 as meaningful (Funder and Ozer said .05 which isn't much better). I mean, come on. At what point do we admit we're scraping the bottom of the barrel here with effect sizes that are noise...and in doing so, misleading the public, policy makers, etc...basically remaining a pseudoscience.

It's just my casual observation, but it seems these threads are correlated...a push for replicability was met by a gross relaxation in effect size interpretation such that it is now meaningless, and we're just doing naïve interpretations of p-values (and hence the larger sample sizes to accommodate this).

Note: I don't think the open science movement was at fault for this...I think this is pure human nature and people knowing which side the bread is buttered on...but I'm less optimistic we're *that* much better off in terms of the product (i.e., good science). But maybe your point is this is incremental...we patched one hole in the sinking ship and that's good even if the ship is still sinking?

I think the large problem is, as the cliché goes, it's hard to get a person to understand something if their job depends upon them not knowing it.

Christopher J Ferguson, Ph.D.'s avatar

Thanks for this comment and it resonated with me.

I think that's been part of what's missing in either "basic" psychology or the open science movement...any philosophical introspection about how well suited our methods are to the thing we are studying.

I still think psychology basically has "physics envy" and very much wants to be a deterministic science. And on some crude levels, sure there's some determinism...genes play an important role in our behavior. A child who is horribly abused is likely to be less well-adjusted than one who is loved and cared for. But beyond some obvious basics, I'm not sure that most of what we do matters even in a small way.

To borrow Lee Jussim's off-the-cuff estimate, probably something like 75% (more maybe) of published psychological science is false positives.

I'm not quite sure what the answer is, but I increasingly wonder if our pseudo-statistical approach to human behavior is much better than opinions with numbers.

Jared Peterson's avatar

Agree on the physics envy. But isnt the demand for increased rigor also a type of physics envy? Is it possible to get precision and control from something as amorphous as a human?

I do research and training for various high stakes domains (eg petrochemical safety, military, law enforcement, medicine) where if our methods are insufficient, people will die. This is a higher bar than most psych research. But mostly our methods are qualitative because controlling for context would defeat the purpose of trying to understand humans who are experts precisely because of how they respond and adapt to context.

Weiss and Shanteau have a paper where they complain about the futility of their research on decision making in real world settings. And I agree their lab based research is irrelevant. But I wonder if the problem is that lab research just isnt a very good paradigm for understanding something as context sensitive as a human? And maybe all the controls and rigor are actually preventing researchers from seeing how humans actually are?

That's how it looks from my side of things. Most quant research just seems so...ephemeral. I don't even know why I should care if it replicates

Christopher J Ferguson, Ph.D.'s avatar

Sadly, after 20+ years of doing this...I increasingly agree. Don't get me wrong, I do think there are a few gems out there.

But I'm increasingly skeptical that *most* quant research tells us anything particularly useful about actual people.

Max Soong's avatar

> I'm still unsure whether the issue is that our methods are not rigorous enough, or whether the problem is the premise that we are discovering hidden truths about human nature that are replicable.

Do you know which side the general consensus is leaning towards, or is the field still at an impasse? I had the (unfortunate) pleasure of choosing a replication crisis topic for my honors thesis which made me believe it was more of the latter, but I'm curious what the current thinking is like.

Christopher J Ferguson, Ph.D.'s avatar

I don’t think there’s anything like a consensus. I think most people prefer not to think of it at all.

Jared Peterson's avatar

Perhaps no consensus, but I would say the dominant narrative in academia (among psychologists) is that the problem is insufficient rigor and perverse incentives. But there is a lot more talk about the need for a paradigm change than there was a decade ago - perhaps in part as a response to the seeming insufficiency of simply increasing rigor.

Darby Saxbe's avatar

I agree, there’s been progress but I’m surprised that NIH/NSF didn’t just require preregistration of all funded projects years ago.

John Wittenbraker's avatar

Great post, especially the Lebowski reference at the end.

Michael Inzlicht's avatar

If you're a fellow Little Urban Achiever, this is the right place for you!

John Wittenbraker's avatar

I’m not sure that I am, but I am good…and thorough.

Michael Inzlicht's avatar

So a fellow brother shamus.

Laurentiu Lupu MD's avatar

Procedural reform can move forward without settling what happens to the old literature. Preregistering the next paper does not by itself change what anyone says about the last one. A field can adopt new methods while leaving its account of the earlier work unresolved.

From the referee's chair that year, did you ever see someone move past a dispute over one finding and start describing their own earlier work differently?

Michael Inzlicht's avatar

Yes there were examples of being reckoning with past work, myself included. Here is a blogpost from that period: https://michaelinzlicht.com/getting-better/2016/2/29/reckoning-with-the-past

Laurentiu Lupu MD's avatar

I read it. The distinction I hadn't made is that the doubt is public before there is a replacement account. You leave the registered replication for a later post, but the uncertainty about your own work is already stated. That gets closer to what I was trying to ask.

Darby Saxbe's avatar

I remember PsychMAP! I think that by the time I checked it out, the period of high drama had mostly ended.

Cindy Hardy's avatar

Thanks for sharing your memories! That debate was a big deal for the discipline of psychology. I hope you now see the positive impact you had on how psychology research methods and statistics are taught and used. I know it caused a lot of discomfort to some in the department I was in, while others tackled the substantive concerns with gusto.

The dude abides!

Dr Lawrence Patihis, PhD's avatar

Sounds like there were problems on both sides, to be honest, including the narcissism of youth and the narcissism of older tenured profs as well.