This study gave people access to an LLM that the researchers knew would give incorrect answers to certain questions, and then quizzed people on those questions, with the option to not respond to a given question if they are unsure about the answer.
This is akin to giving someone a textbook on an obscure subject that has certain factual errors, letting them know they can use that textbook in a quiz on that subject, and then quizzing that person on those facts that the textbook gets wrong.
Obviously that person is both more likely to be willing to respond to the question and is more likely to get it wrong!
There are a lot of things I'm very interested in that are specific to modern LLMs and how they affect learning and confidence (sycophancy, cognitive helplessness, etc.).
This study tested none of those. Its experimental setup is not very different than simply substituting the LLM with a textbook with errors.
Even if AI gets smarter it will still be agreeable, and people will use it to more confidently reinforce thier stupider ideas, especially in areas where they lack the knowledge to know if they are right.
Even if this can be solved, technically, people won't want to use the model that says they are wrong, so they will choose the glib lies that reinforce their beliefs.
Freedown of speech forces people to think they may be wrong, but freedom of association lets people avoid this. AI is going to be internet hug box echo chambers at an unbelievable scale.
Richard Feynman: "I have the advantage of having found out how hard it is to get to really know something, how careful you have to be about checking the experiments, how easy it is to make mistakes and fool yourself..."
Id want to know if “AI” makes a material difference vs just having access to the wrong answer. Like someone could be given search access that successfully retrieved wrong answers to questions, would that give the same results. How much do uniquely AI characteristics, like sycophancy or the conversational aspect play into this, vs people just being willing to believe what they read?
$0.10 awarded for a correct answer
$0.10 deducted for a wrong answer
$0.00 (no change for a refusal to answer)
There was no detail on whether participants were actually going to be paid, or if they were in the negative at the end, be expected to pay their losses.I am inclined to believe the effects of the study are real, but not nearly as pronounced as the data. If there were more serious amounts of money on the table, I think common sense would prevail.
I feel like this might be an interesting method for places where the default response tends to be someone submitting copy-pasta: just give the ai-response and then have community discussion around what it missed.
You can’t say it didn’t warn us, it’s on the tin: “AI might be wrong.”
I get why they used questions where AI models fail, but it also really reduces the value of this study. Nobody is really asking AI the color of a team's uniform in a movie, and if they do and confidently get it wrong, it just doesn't matter at all.
Asking trivial questions also feels like it would affect the rate at which people are willing to confidently say things that are wrong. If you ask me some question of pointless trivia and I ask ChatGPT, I'll probably just repeat the answer because who cares. If you ask me something even mildly important and I ask ChatGPT, I'll either verify the information before I repeat it to you, or I'll qualify that I looked it up with ChatGPT and didn't verify. But some things are just so unimportant that they don't even warrant the disclaimer.
> Researchers found AI advice suppressed judgment suspension from 44% to 3%, accuracy from 27% to 9%, while confidence rose from 30% to 76%. People trusted wrong AI answers.
I didn't see the article mentioned what they were actually asked, but I'm surprised confidence was only ~30%.
And people who follow bad advice will get bad results, and people who follow good advice will get good results?
I wonder if this will impact the quality and prevalence of certain AI models in the future.
> sudy
Intentional?
https://www.youtube.com/watch?v=axOcn--n_lM
https://www.anthropic.com/research/AI-assistance-coding-skil...
We should remember LLM do have legitimate use-cases like search, as we enter the "Trough of disillusionment" in the hype cycle. =3
There's plenty of other issues with the study, one of the primary ones being that they chose a relatively 'dumb' AI (Step 3.5 Flash), but also, specifically hand-selected wrong answers. People that are used to using competent AIs that are mostly correct would be operating off of experience that suggests they should trust AI. In this case, by design, they shouldn't have, but it's hard to fault the user for that.