🧭 What decades of research on trust in automation say about working with AI today - and the habits it changed in how I design the systems I work on.

📧 Found an error or have a question? write to me or leave a comment

A grey dome tent on dry grass beside a wooden boardwalk that curves away into thick fog over the Sibillini mountains

Fog rolling in over the Sibillini: a well-built path, and no way to see where it ends. | Image by Author


I spent the summer reading papers on trust in automation, some of them more than twenty years old. The underlying question looks technical: when to trust a machine, and when to stop. I didn’t find a single rule. I did find quite a few things that surprised me, and I tried them out on the projects I work on. Some held up; others shrank as I went.

🎚️ Calibrated trust, not high trust Link to heading

For years I thought the problem was getting users to trust the system more. That was my first surprise. I was wrong, and for a structural reason. The cost of overtrust is invisible and you pay it late: until the error nobody checked shows up, everything looks fine. The cost of underuse is a red KPI that management sees the week after release. So people who work on “increasing trust” end up treating the problem they can see right away, low adoption, and neglecting the one that will surface months later, the errors nobody checked.

Lee and See made the case for calibrated trust back in 2004 [1]. Trust in an automated system rests on three bases: performance, meaning what it does and how well; process, how it gets there; and purpose, why it exists and what it was built for. With an LLM, process is the hardest basis to provide honestly. A reasoning trace looks like an account of how the model reached its answer, but nothing guarantees it reflects the actual computation, and the studies that measured it found it often doesn’t [2]. Showing it risks miscalibration dressed up as calibration, which is worse than saying nothing: users believe they’ve understood, when all they’ve read is a plausible story.

Calibrated trustTrust plotted against what the system can actually do: trust above the calibration line is overtrust, whose cost is invisible and paid late; below it is distrust, whose cost is visible and paid right away. The reviewer of the extraction pipeline sits in the overtrust area, trusting at the level first promised rather than the level the system delivers.TrustCapabilityCalibrated trustOvertrust → misusecost: invisible, paid lateDistrust → disusecost: visible, paid nowThe reviewer’s trustheld where the first promise put itthe gap: errors nobody checks
Calibrated trustTrust plotted against what the system can actually do: trust above the calibration line is overtrust, whose cost is invisible and paid late; below it is distrust, whose cost is visible and paid right away. The reviewer of the extraction pipeline sits in the overtrust area, trusting at the level first promised rather than the level the system delivers.TrustCapabilityCalibrated trustOvertrust → misuseinvisible cost,paid lateDistrust → disusevisible cost, paid nowThe reviewer’s trustheld where the firstpromise put itthe gap: errorsnobody checks

Calibrated trust, after Lee and See: trust should track what the system actually does. The dot is the reviewer from the extraction pipeline below. | Diagram by Author

I saw this on a document extraction system, a PDF pipeline with a human reviewer downstream. Once the reviewer’s trust settles, the share of fields corrected by hand stops measuring the system’s error rate and starts measuring how much the reviewer actually checks. Then there was the purpose inherited from the original project, which promised to replace the reviewer, while today the system assists them. Users trusted it at the level promised at the start; the system worked at a different one. I still ask users how much they trust the system, but the answer on its own tells me little. What I try to do is track and collect as many real numbers as I can: how often the system is actually wrong in the field, and how much people actually check. Only by putting the two side by side can I tell whether trust sits at the right level. And when someone on a project suggests showing the model’s reasoning as a transparency guarantee, I make it clear it’s no cure-all. It can be useful, but it doesn’t really tell you how the model got to its answer.

🧑‍🔧 Abuse is the designer’s problem, not the user’s Link to heading

When automation gets rejected or misused, the temptation is to blame the people using it: they trust it too much, or too little. In 1997 Parasuraman and Riley [3] named four ways people relate to automation. There’s use, deciding to hand it a task; misuse, relying on it too much and no longer checking it; disuse, ignoring it or switching it off, often after too many false alarms. And there’s abuse, which has nothing to do with the operator: automating without considering the consequences for human performance and for the authority of the people who work with the system. The first three come up all the time, the fourth almost never. The paper has a line I still paste into reports: to the extent that a system is made less vulnerable to operator error, it is made more vulnerable to designer error. Human error doesn’t leave the equation. It moves upstream, where it’s harder to see.

Use, misuse, disuse, abuseParasuraman and Riley's four relations with automation: use, misuse and disuse belong to the operator, while abuse belongs to the designer, and making a system less vulnerable to operator error shifts the error upstream to the designer.The designerThe operatorError shiftsAbuseautomating withoutweighing the cost forthe people who use itUsehand ita taskMisuserely on it,stop checkingDisuseignore it,switch it off“To the extent that a system is made less vulnerable to operator error,it is made more vulnerable to designer error.”
Use, misuse, disuse, abuseParasuraman and Riley's four relations with automation: use, misuse and disuse belong to the operator, while abuse belongs to the designer, and making a system less vulnerable to operator error shifts the error upstream to the designer.The designerThe operatorError shiftsAbuseautomating without weighing the costfor the people who use itUsehand ita taskMisuserely on it,stop checkingDisuseignore it,switch it off“To the extent that a system is made lessvulnerable to operator error, it is mademore vulnerable to designer error.”

Parasuraman and Riley’s four relations with automation: three belong to the operator, one to the designer. | Diagram by Author

Applied to the question in the title, this means that before asking whether the operator trusts the system the right way, you have to ask whether the system, as designed, deserved that trust. The operator’s role should be defined around their actual responsibilities and abilities instead of falling out as a by-product of the implementation. Otherwise you plant the problem before you’ve even shipped.

I watched this happen with an automation that management decided on without involving the people who would use it every day. The friction that followed was read internally as resistance to change. Looked at more closely, it was abuse upstream, and the operators were flagging a design problem in the only way they had. Over the last few months I’ve applied this lens to three real cases, and in all three, what looked like an adoption problem turned out, once I dug in, to be a decision made without the people who had to live with the result. Three cases aren’t statistics, but they were enough to change a habit. In scoping, my first question used to be about requirements. Now it’s who, among the people who will use the system every day, was in the room when the decision to build it was made. If the answer is nobody, I already know where to look when something goes wrong.

I do keep a counterweight, though, so I don’t end up putting anyone’s intentions on trial. It’s my own, not the paper’s. Even a heavy degradation can pay off in the long run: the arrival of PCs probably brought months, maybe years, of lower productivity compared with pen and paper. Not every first-quarter adoption failure is a design error; sometimes it’s just the cost of the transition.

🧠 The interventions that work are the ones people hate Link to heading

The result that surprised me most comes from Buçinca, Malaya and Gajos [4], who in 2021 tested cognitive forcing functions (CFFs), interface interventions designed to make users reason through a decision instead of accepting the AI’s output. They tried three. On Demand shows the AI’s answer only if the user asks for it; Update asks users to decide on their own before seeing it, then lets them change their mind; Wait shows it only after a forced delay. The interventions that cut overreliance the most are exactly the ones users like least, and the benefit only shows up when the AI is wrong. On a system that is almost always right, trusting it blindly still gets you close to optimal results, so you pay the cost of the intervention even when you didn’t need it.

The mechanism is the familiar split between fast and slow thinking [5]. Instead of evaluating each output, people build a general heuristic about how far to trust a system, and CFFs break that shortcut by forcing a slower evaluation. The more an intervention makes people think, the better it works and the less they like it.

The paper has a second, less comfortable result, which the authors found by auditing the inequalities their own intervention produced. The benefit isn’t evenly spread. People high in need for cognition, the trait that measures how motivated someone is to stop and think things through, gain the most. People low in it improve only partly, and not significantly on overall performance. It has little to do with intelligence, since the literature finds only a modest correlation with cognitive ability [6]; it’s mostly a willingness to think. The people who most need the push, in other words, are the ones who get the least out of it. In presales I don’t use this to profile anyone. I use it as a design question: for which users of this system will a CFF be a tool, and for which will it just be an obstacle to work around?

Two mechanisms, and one idea I dropped Link to heading

I tried applying this logic to a real multi-agent system, with two mechanisms. I called the first one ghost track. Before execution, the operator sketches a high-level plan using the same action vocabulary as the agent, and if execution diverges structurally from that plan, a lightweight gate asks whether to continue or stop. Having the AI propose the plan was tempting, but the plan has to stay the human’s, otherwise you lose the very independence that makes the comparison useful. The second is a mission-critical interrogation: before a critical action the agent states what it’s about to do, the operator argues the pros and cons, and a second model rates the quality of that argument. In both cases I accepted some residual misuse to keep the gate lightweight, which is exactly the trade-off the paper describes.

Ghost trackThe operator sketches a plan before the run in the agent's own action vocabulary; the agent's execution is compared with it, and only a structural divergence triggers a lightweight gate asking whether to continue or stop, otherwise the run continues.YesNoOperator’s plansketched before runAgent’s executionwhat actually runsStructuraldivergence?Lightweight gateoperator decides: continue or stopRun continuesThe plan has to stay human:an AI-drafted one would comparethe agent with itself.
Ghost trackThe operator sketches a plan before the run in the agent's own action vocabulary; the agent's execution is compared with it, and only a structural divergence triggers a lightweight gate asking whether to continue or stop, otherwise the run continues.NoYesOperator’s plansketched before runAgent’s executionwhat actually runsStructuraldivergence?Run continuesLightweight gateoperator decides: continue or stopThe plan has to stay human: an AI-draftedone would compare the agent with itself.

Ghost track: the operator’s plan runs alongside the agent, and only a structural divergence stops the run. | Diagram by Author

Mission-critical interrogationBefore a critical action the agent states what it is about to do, the operator argues both pros and cons, a second model rates the quality of that argument, and only then does the operator let the action go ahead or stop it.1 · states action2 · pros and cons3 · rates argument4 · go or stopAgentOperatorSecond model
Mission-critical interrogationBefore a critical action the agent states what it is about to do, the operator argues both pros and cons, a second model rates the quality of that argument, and only then does the operator let the action go ahead or stop it.1 · states action2 · pros and cons3 · rates argument4 · go or stopAgentOperatorSecondmodel

Mission-critical interrogation: the operator argues both sides before a critical action, and a second model rates the argument. | Diagram by Author

One idea I dropped precisely because of the paper. Letting the user pick between options A, B or C to steer the machine’s reasoning seemed elegant, until I realized that from the user’s side it was still the experiment’s baseline condition: they see the AI’s proposal and its reasoning, and they evaluate them. That condition doesn’t reduce overreliance. The anchoring risk stays the same and only the packaging changes. Without the paper I’d have built it anyway and found the problem in production.

🔍 Explanations don’t save the human-AI team Link to heading

Accuracy isn’t the metric Link to heading

Bansal and colleagues’ work on explanations [7] hurt me as an engineer, because it takes apart an intuition I took for granted for years: that showing users why the system answered the way it did helps them trust it better. They measured it, comparing explanations with simply showing the model’s confidence, and the net effect is close to zero. Explanations raise team accuracy when the AI is right and lower it when it’s wrong, and the two effects nearly cancel out. Even explanations written by hand by experts do no better than generated ones.

What explanations do to team accuracyIllustrative view of Bansal and colleagues' result: explanations raise human-AI team accuracy when the AI is right and lower it when the AI is wrong, so the net effect is close to zero.Worse teamBetter teamWhen the AI is rightthey helpWhen the AI is wrongthey hurtNet effectclose to zero
What explanations do to team accuracyIllustrative view of Bansal and colleagues' result: explanations raise human-AI team accuracy when the AI is right and lower it when the AI is wrong, so the net effect is close to zero.Worse teamBetter teamWhen the AI is rightthey helpWhen the AI is wrongthey hurtNet effectclose to zero

What explanations did to team accuracy in Bansal and colleagues’ study, illustrative, not to scale. | Diagram by Author

The useful question, then, is whether the team does better than either the AI or the human alone, what the paper calls complementary performance. How accurate each one is on its own doesn’t answer it. A more accurate AI, paired with convincing explanations, can make the team worse in exactly the cases that matter, the ones where it’s wrong.

The most convincing account of this result I found comes from a 1991 paper. Koehler [8] shows that treating a hypothesis as true, even temporarily in order to evaluate it, makes it seem more plausible, regardless of how good the argument is. You don’t need a good explanation. The task of taking it seriously is enough to make it more believable than it deserves.

On one point the paper left me halfway. In the multi-agent system I work on, human-only performance doesn’t exist, because nobody does that task alone, so there’s no baseline for saying whether the team is complementary. It’s no accident that the paper deliberately picks tasks where the human and the AI, each on their own, perform about the same. The paper’s framework doesn’t apply as is, and the right metric still has to be built.

What changed in my work Link to heading

Day to day, this has turned into a few new habits. When a client asks how sure the system is of an answer, before replying I ask whether there’s an independent way to measure its accuracy. Without one, any confidence number risks being another reasoning trace, plausible and unverified. For generative outputs there are approaches like semantic entropy [9], which samples the same question several times and measures how far the answers agree in meaning. It isn’t an independent measure of accuracy, since a model that is always wrong in the same way converges nicely, but it does catch answers made up on the spot. I’d got there on my own before finding out it was already a research line, which tells me the problem is still open and more than an implementation detail. I’ve also grown wary of prompts that ask the model to argue for an answer before giving it, because arguing alone makes the answer more convincing. The same effect made me revisit my own design: asking the operator to argue about the agent’s action risks making that action look more plausible to them. That’s why the interrogation asks for pros and cons, since forcing yourself to consider the alternative is the classic countermeasure.

👀 Severity depends on who’s looking Link to heading

The last surprise came from an audit I did myself, not from a paper read on the couch, and it concerns who decides when a risk is serious enough to deserve less trust. Amershi and colleagues [10] compiled 18 guidelines for human-AI interaction, now part of Microsoft’s HAX Toolkit, grouped by moment: initially, during interaction, when the system is wrong, and over time. Applying them to a tool in production, I found seven concrete violations and three guidelines already met; the other eight didn’t apply to that kind of tool.

I called one of the violations “silent confirmation”. A section no human has ever opened is marked as confirmed by default, and the system shows “8 of 8 confirmed” before a single real review has happened. For someone who studies trust calibration this is a high risk, because it enables overreliance in its purest form. For the people who use the product every day the same pattern is neutral, if not a convenience: ok, ok, export, done. The design the guidelines call correct doesn’t necessarily match what the market rewards, and a count of violations says little unless you ask compared to what, and seen by whom.

What surprised me more than the violation was my own reaction. The silent confirmation bothered me much sooner and much more than anyone else in the room, including the person who uses that product every day, and I spent more time asking myself why it bothered me so much than explaining it to the people who had to decide. The tension between research and market, it turns out, runs through me as well as through clients: it’s one of my biases too. Now, before I raise a red flag on a guideline, I ask myself whether I’m protecting the user or defending a principle that only seems obvious to me.

🧭 The question I’m left with Link to heading

There’s a tacit assumption I only noticed when I reread these papers together instead of one at a time: when automation meets something it has never seen, it fails, and the human takes over. For a 1997 autopilot that holds, because when it leaves its envelope it disengages and everyone notices. An LLM agent is built to generalize. Faced with something new it doesn’t stop; it still produces a plausible answer. The failure, if there is one, goes silent, and we’re back to the cost of overtrust that nobody sees.

Off its own ground: autopilot vs LLM agentFaced with a situation it has never seen, a 1997 autopilot leaves its envelope and disengages so everyone notices, while an LLM agent generalizes and still produces a plausible answer, so the failure goes silent.Autopilot, 1997LLM agent, todaySomething newnever seen beforeLeaves its envelopeand disengagesLoud failureeveryone noticesSomething newnever seen beforeGeneralizesit doesn’t stopPlausible answerthe failure, if any, goes silent
Off its own ground: autopilot vs LLM agentFaced with a situation it has never seen, a 1997 autopilot leaves its envelope and disengages so everyone notices, while an LLM agent generalizes and still produces a plausible answer, so the failure goes silent.Autopilot, 1997LLM agent, todaySomething newnever seen beforeLeaves its envelopeand disengagesLoud failureeveryone noticesSomething newnever seen beforeGeneralizesit doesn’t stopPlausible answerfailure goes silent

The tacit assumption: automation that meets something new fails loudly. An LLM agent doesn’t. | Diagram by Author

How do you calibrate trust in a machine that doesn’t tell you when it’s off its own ground? I don’t have an answer yet. I’ll come back to it.


🔗 Sources Link to heading

[1] Lee, J.D. & See, K.A. (2004), Trust in Automation: Designing for Appropriate Reliance, Human Factors 46(1), 50-80. SAGE

[2] Turpin, M., Michael, J., Perez, E. & Bowman, S.R. (2023), Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, NeurIPS 2023.

[3] Parasuraman, R. & Riley, V. (1997), Humans and Automation: Use, Misuse, Disuse, Abuse, Human Factors 39(2), 230-253. SAGE

[4] Buçinca, Z., Malaya, M.B. & Gajos, K.Z. (2021), To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making, Proc. ACM HCI 5, CSCW1. PDF

[5] Kahneman, D. (2011), Thinking, Fast and Slow, Farrar, Straus and Giroux.

[6] Cacioppo, J.T., Petty, R.E., Feinstein, J.A. & Jarvis, W.B.G. (1996), Dispositional differences in cognitive motivation: The life and times of individuals varying in need for cognition, Psychological Bulletin 119(2), 197-253.

[7] Bansal, G. et al. (2021), Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance, CHI ‘21. arXiv PDF

[8] Koehler, D.J. (1991), Explanation, imagination, and confidence in judgment, Psychological Bulletin 110(3), 499-519.

[9] Farquhar, S., Kossen, J., Kuhn, L. & Gal, Y. (2024), Detecting hallucinations in large language models using semantic entropy, Nature 630, 625-630.

[10] Amershi, S. et al. (2019), Guidelines for Human-AI Interaction, CHI ‘19. PDF · HAX Toolkit

[11] Parasuraman, R., Sheridan, T.B. & Wickens, C.D. (2000), A model for types and levels of human interaction with automation, IEEE Transactions on Systems, Man, and Cybernetics, Part A 30(3), 286-297.

[12] Klein, G., Woods, D.D., Bradshaw, J.M., Hoffman, R.R. & Feltovich, P.J. (2004), Ten Challenges for Making Automation a “Team Player” in Joint Human-Agent Activity, IEEE Intelligent Systems 19(6), 91-95.