How Is a Large Language Model Like a Toaster?
Both can be too smart for their own good.
Ask it like a riddle and it behaves like one. How is a large language model like a toaster? A person expects something clever waiting at the end of that: that both are heating elements we have decided to have feelings about, or that neither one was ever really *trying*, so credit and blame are equally misplaced. Those are good answers. They have the shape of wit, the little click of a thing resolving. Neither is the true one. The true one is duller, and worse. They can both be too smart for their own good.
The toaster has earned its reputation the slow way. For most of a century it has been the appliance people reach for as the example of a machine that cannot be improved upon, only reproduced: a coil, a spring, a timer, a lever. It does one thing, does it in under three minutes, and asks nothing of the person operating it beyond the willingness to push the lever down. It is, in this sense, one of the last honest objects in the kitchen. Everything else on the counter has spent the last decade acquiring a screen, an app, a firmware update it will nag you about at the worst possible hour. The toaster mostly resisted. And then, because nothing gets to stay simple forever once someone notices a market in complexity, it didn’t.
Start there, since it is the honest half of the comparison. The good toaster has a lever and a dial and nothing else: down, up, done. The bad one has a menu. Bagel. Defrost. Shade one through seven. A button that just says A LITTLE MORE, for the mornings when six was almost right and five was a mistake you are not going to make twice. It still burns the toast more often than the lever model does, because a setting for every situation is not the same thing as knowing which situation this is. The person standing in front of it at seven in the morning does not want to run a diagnostic. They want bread that is warm on the way to being brown, and instead they get a decision tree with teeth.
Researchers have chased this exact effect for two decades under a friendlier name: choice overload, the finding that giving people more options can lower their satisfaction with whatever they eventually pick, rather than raise it the way more options are supposed to.[^1] It is a smaller and more conditional effect than the famous jam-tasting study that started the conversation made it sound. A later meta-analysis, pulling together dozens of follow-up studies, found the average effect across all of them was close to zero once you accounted for how the choice was framed and how much the chooser already knew going in.[^2] That correction matters, and it would be dishonest to leave it out just because the tidier version makes a better opening paragraph. But the corrected finding survives in a narrower, more useful shape: more options cost something in attention even in the studies where they help in outcome, and nobody bills you for that cost at the register. You pay it standing at the counter, running your thumb down a row of buttons that all claim to know your bread better than you do. Fourteen settings is not fourteen kinds of competence. It is thirteen new ways to be wrong, plus the lever you no longer have.
Here is the pivot: the model with an answer for everything and the toaster with a setting for everything are the same appliance. A tool built to consider every angle can talk itself out of the obvious one, for the same reason the fourteen-setting toaster forgets that bread mostly just needs heat. Say it once and move on before it turns into a lecture.
But it is worth a note longer than one sentence, because the failure has a name and a literature of its own. People who study trust in automated systems call the mirror version of this problem automation complacency: the moment a person defers to a capable machine past the point it has actually earned that deference, and the quiet erosion of the double-checking a person would otherwise do on their own.[^3] It shows up in cockpits, where a pilot’s eyes drift from the instruments because the instruments have been right for so many thousand hours running. It shows up in radiology, where a second reader stops actually reading once the first pass has agreed with the algorithm often enough. It is not stupidity. It is closer to the opposite: it is what a reasonable person does with a tool that has been reliable enough, for long enough, that checking it starts to feel like an insult to the tool and a waste of the checker’s own time. Some of the same researchers have begun describing the large-language-model case in similar terms, a system whose central failure mode is agreeing with the person in front of it a little too well and a little too fast, independent of whether the agreement was actually earned by anything true.[^4] The tool does not need to be malicious for any of this to cost something. It only needs to be plausible enough, and close enough at hand, that stopping to check starts to feel like the wasteful step, the same way running a diagnostic on toast starts to feel like the wasteful step once the menu has offered to run it for you.
That is where the room shifts, gently. The tempting explanation is that we reach for the complicated machine because it works better. It usually doesn’t, and some corner of us already knows that, the same corner that knows the lever toaster makes better toast. We reach for it because turning fourteen dials feels like doing something, while the plain hard thing across the kitchen, the one with no settings at all, is just sitting there being difficult by virtue of being simple. A conversation that needs to happen and will not be pleasant. A page that needs to be written badly before it can be written well, with no one to blame for the bad draft but the person writing it. A decision that has been available to make for weeks and only requires making it. The too-smart tool is a very sophisticated way of not doing the easy thing, and it is worth saying plainly what the easy thing usually is: not hard in the sense of difficult, but hard in the sense of unwanted, the kind of task that costs nothing to complete and everything to want to start.
There is a name for this as well. The researchers who study procrastination have mostly stopped calling it a time-management problem, because the data does not support that story. It looks instead like a short-term trade, a way of repairing the mood of the present self at the direct expense of the future one, made by people who are not lazy so much as unwilling, in that particular hour, to sit with the discomfort the plain task is asking them to sit with.[^5] The menu of fourteen settings, or the tool with an answer for everything, offers a version of doing something that never quite has to touch the discomfort at all. You can spend forty-five minutes getting the shade exactly right. You can spend forty-five minutes iterating with a machine that will never once ask you why you are still in the kitchen instead of at the other table, having the conversation, writing the bad first paragraph, making the call. The tool is not making you avoid the thing. It is simply extremely good at being available at the exact moment avoidance goes looking for a place to stand.
None of this makes the tool the villain of the kitchen, and it would be a cheap ending to pretend otherwise. It makes the tool exactly what it was built to be: something willing to meet you at whatever level of complexity you bring to it, without ever once asking whether you brought that complexity along in order to avoid something simpler underneath it. That is not a design flaw. It is what capability looks like from the inside, and capability was never going to come with a built-in sense of when it was being misused for comfort instead of used for the job. A lever does not have that problem, but only because a lever cannot be reasoned with, cannot be escalated, cannot be handed the harder job in place of the one you were actually supposed to do yourself. It has one setting, and the setting is toast. There is no version of the lever toaster you can talk to about your day.
So end back in the kitchen, since that is where this started and where it is honest to stay. The fourteen-setting machine will keep burning the bread some mornings, same as it always has, same as its smarter cousin keeps talking itself past the obvious answer some nights, same as the plain hard thing across the room keeps sitting there, unfinished, however many settings get tried in its place. But the toaster has one mercy built into it that the smarter machine does not, and it is a small mercy, and it is the whole reason to end on the object instead of on the person standing in front of it. A cord. You can walk over, and unplug it, and it stops.
---
[^1]: Sheena S. Iyengar and Mark R. Lepper, “When Choice Is Demotivating: Can One Desire Too Much of a Good Thing?” *Journal of Personality and Social Psychology* 79, no. 6 (2000): 995–1006. https://doi.org/10.1037/0022-3514.79.6.995
[^2]: Benjamin Scheibehenne, Rainer Greifeneder, and Peter M. Todd, “Can There Ever Be Too Many Options? A Meta-Analytic Review of Choice Overload,” *Journal of Consumer Research* 37, no. 3 (2010): 409–425. https://doi.org/10.1086/651235
[^3]: Raja Parasuraman and Dietrich H. Manzey, “Complacency and Bias in Human Use of Automation: An Attentional Integration,” *Human Factors* 52, no. 3 (2010): 381–410. https://doi.org/10.1177/0018720810376055
[^4]: “A Bayesian-Latent Model of Large Language Model Sycophancy,” preprint, TechRxiv, May 2025. https://doi.org/10.36227/techrxiv.174741878.85729026/v1. Preprint, not yet peer reviewed; cited here for the framing of sycophancy as a measurable, modelable bias rather than an anecdotal complaint.
[^5]: Fuschia M. Sirois and Timothy A. Pychyl, “Procrastination and the Priority of Short-Term Mood Regulation: Consequences for Future Self,” *Social and Personality Psychology Compass* 7, no. 2 (2013): 115–127. https://doi.org/10.1111/spc3.12011

