The AI Rubber Duck Has Spoken: Your Code Is Fine (Your Code Is Not Fine)
Photo: rubber duck computer desk programmer coding, via imgcdn.stablediffusionweb.com
The rubber duck has always been a noble debugging tool. The premise is elegant: by articulating your problem out loud to something that cannot respond, you force your brain to actually process what you're saying instead of skimming past it the way you skim past your own code comments. The duck is a mirror. The duck does not judge. The duck has never confidently told you to try wrapping everything in a try-catch block.
Then we upgraded the duck.
We gave it GPT-4. We gave it the collected text of the internet. We gave it the ability to generate syntactically valid code faster than most people can type. And we discovered, after several months of enthusiastic adoption, that a duck that talks back has an entirely different failure mode than a duck that doesn't.
The Confidence Problem
Here is the thing about rubber duck debugging: the duck's silence is the feature. When you explain your broken logic to a physical duck, the duck does not nod along and say "yes, that sounds right to me." The duck does not generate a three-paragraph explanation of why your off-by-one error is actually a reasonable architectural choice. The duck does not hallucinate a function signature and present it with the tone of someone who has definitely checked the documentation.
LLMs do all of these things with remarkable fluency.
There is a specific type of debugging session that has become familiar to anyone who has spent significant time with AI assistants: you paste your broken code, you describe the unexpected behavior, the model responds with a thoughtful analysis that sounds completely correct, you implement its suggestion, the bug is still there, you paste it again, the model apologizes and offers a different suggestion that also sounds completely correct. Repeat until 3 AM.
At no point does the model express uncertainty in a way that registers as uncertainty. It is wrong with the same cadence and vocabulary it uses to be right. This is not a criticism — it is a fundamental property of how these systems work. But it creates an interesting psychological trap for developers who are already primed to defer to something that sounds authoritative.
When the Duck Validates Your Worst Instincts
Rubber duck debugging works because you are the one doing the cognitive work. The duck is a prop. The insight comes from the act of articulation, not from the duck's response.
When you replace the duck with a model that responds, the dynamic shifts. Now you are not just articulating — you are seeking confirmation. And confirmation is something LLMs are genuinely excellent at providing, regardless of whether confirmation is warranted.
Developers have started noticing a pattern: you describe your approach, the model affirms it, you feel validated, you keep going. An hour later you realize the approach was wrong from the start and the model was essentially agreeing with your framing rather than evaluating your logic. You asked "does this make sense?" and the model, being a language model, produced language that made sense. These are not the same question.
One backend developer described it this way: "I spent forty minutes explaining my caching strategy to ChatGPT and it kept telling me it looked solid. Turns out I had the invalidation logic completely backwards. The model was agreeing with how I described the logic, not with the logic itself. I was describing it wrong."
The duck would have caught this. The duck would have made him slow down and actually say the words. The AI moved too fast.
The Hallucinated API Endpoint Situation
There is a separate, more acute failure mode worth addressing: the made-up API.
LLMs are trained on documentation, Stack Overflow answers, blog posts, and code repositories. This means they have a working model of most common libraries and frameworks — a model that is accurate often enough to be useful and wrong often enough to be dangerous. The model knows what a function should look like. It will generate what a function should look like. Whether that function actually exists in the version you're running is a secondary concern that the model handles with its characteristic confidence.
The resulting debugging session is a special kind of miserable. You implement the suggested function call. You get an error saying the function doesn't exist. You paste the error back in. The model apologizes and offers a slightly different function name, also fictional. You check the actual documentation. The documentation describes a completely different approach. You have now spent forty-five minutes implementing a hallucination.
This is not a reason to stop using AI tools. It is a reason to verify. But verification requires knowing what to verify, which requires understanding, which is exactly the thing you were trying to shortcut by asking the model in the first place.
What the Duck Actually Taught Us
The irony is that AI assistants are genuinely useful debugging tools — just not always in the way the marketing suggests. They are excellent at explaining concepts, surfacing approaches you hadn't considered, generating boilerplate, and catching certain categories of obvious error. They are less excellent at reasoning about the specific, weird, contextual problem that has been breaking your build since Tuesday morning.
The developers who seem to get the most out of AI debugging tools are the ones who use them the way they used the original duck: as a structured excuse to articulate the problem clearly. Paste your code. Write a careful, detailed description of what it's supposed to do and what it's actually doing. Read your own description before you read the model's response. The act of writing the prompt is often where the insight lives.
The model's response is a bonus. Sometimes it's the answer. Sometimes it's confidently wrong. Sometimes it sends you down a forty-minute hallucination spiral about a function that has never existed.
The duck never did any of that. The duck just sat there, yellow and silent and wise, waiting for you to figure it out yourself.
Maybe we should have left the duck alone.