ChatGPT cons and limits, and how to work around them

Jack 19 JULY 2026 10 min read

The cons of ChatGPT that actually cost a business aren’t the obvious ones. It isn’t that it can’t draw well or that the free plan runs out mid-afternoon. It’s that it will hand you a wrong answer in the same confident voice it uses for a right one, and if you act on it, the mistake is yours to wear. Everything below is a limit I hit using it every day for real work, where it bites, and what to do so it doesn’t.

It still makes things up, and it never fully stops

ChatGPT invents facts, and no release has cured it, but the pattern isn’t the straight line people assume. On OpenAI’s own PersonQA fact test, its 2025 reasoning model o3 gave a wrong answer 33% of the time against the older o1’s 16%, and the smaller o4-mini missed 48% (TechCrunch’s write-up of the system card): the models built to reason harder guessed more, not less. Then GPT-5 pulled it back hard, with OpenAI reporting its thinking model makes roughly five times fewer factual errors than o3. So it lurches better and worse from one model to the next, but it has never reached zero and it won’t, because the tendency is built into how the thing works, not bolted on by accident.

Why it’s structural, not a bug waiting for a patch: in a September 2025 paper, Why Language Models Hallucinate, OpenAI’s own researchers describe a model trained and graded in a way that rewards a confident guess over “I don’t know”, the same way a student low on time is better off guessing than leaving a blank. Silence scores zero; a plausible guess sometimes scores. So the machine guesses, fluently, every time. If you want the mechanism laid out in a few minutes, IBM’s Why Large Language Models Hallucinate explains why a system that predicts the next word will always be willing to invent one.

For a business that turns into one habit: anything with a fact in it you’d act on, a name, a figure, a citation, a date, gets checked before it leaves your desk.

Wrong and confident is the dangerous combination

A confident wrong answer is more dangerous than a hesitant one, because confidence is how people decide what to trust. A colleague who isn’t sure sounds unsure. ChatGPT sounds identical whether it’s spot on or fabricating, so the tell you normally rely on is gone. That’s why it catches out people who are anything but careless.

In October 2025, Deloitte refunded the final instalment of a A$440,000 (about US$290,000) government contract after a report it delivered to the Department of Employment and Workplace Relations was found to contain academic references that didn’t exist and a quote attributed to a Federal Court judge who never said it, drafted with the help of Azure OpenAI’s GPT-4o (Fortune). In 2024, a British Columbia tribunal held Air Canada liable when the chatbot on its own website invented a bereavement refund policy that wasn’t real; the airline argued the bot was “a separate legal entity responsible for its own actions”, and the tribunal disagreed (CBC). Put an AI in front of customers and its made-up answer becomes your promise to honour.

The one that should really land is the lawyers, because checking citations is their entire job and the tool still walks them into it. Back in 2023 a New York lawyer was fined US$5,000 for filing six court cases ChatGPT had fabricated; when he asked the model whether they were real, it assured him they were (Mata v. Avianca). That was the first to make headlines, not the last. A legal researcher now keeps a public database of court decisions caught citing AI-invented cases, and it passed 1,500 by mid-2026, running at roughly eight new ones a day, with sanctions climbing from that original $5,000 towards $15,000 an attorney and, in some courts, suspension.

None of these people were reckless. They trusted fluent output the way you’d trust a knowledgeable colleague. The tool talks like one and isn’t one.

Fluency is no evidence of truth. ChatGPT sounds exactly as certain when it’s inventing as when it’s right, so its confidence tells you nothing about whether it’s correct. Check anything you’d act on, not because the tool is bad, but because the usual signal for “this is solid” has been switched off.

Its knowledge stops at a date

ChatGPT’s knowledge is frozen at a cutoff, and it’s usually months behind the day you’re using it. Every model is trained up to a point and then sealed. As of 2026 the paid default’s training runs to late 2025, and the free tiers are older still. It can search the web when it chooses to, which patches the gap, but when it answers from memory it’s answering from a snapshot that may predate the thing you asked about.

Ask it today’s price for something, this year’s tax threshold, who currently holds a role, or whether a rule still applies, and a stale answer looks exactly like a current one. The fix is to make the search explicit: tell it to look it up and show you the source, and treat any current-sounding “fact” with no source as a maybe until you’ve confirmed it.

It can’t do maths on its own

ChatGPT can’t be trusted with arithmetic unaided, because it predicts text, it doesn’t calculate. Ask it to total a column of figures or work out a margin and it produces something shaped like maths that is sometimes wrong, and wrong in a way that’s easy to miss because the formatting is flawless. The well-known version is it miscounting the letters in a word. The version that costs you is a quietly incorrect total sitting inside a quote you’re about to send.

Modern ChatGPT often hands the sum to a built-in calculator or writes a scrap of code to work it out, which fixes most of it, but only when it decides to reach for that tool. So for anything numeric that matters, tell it to calculate with its data-analysis tool rather than in its head, or check the figure yourself before it counts.

The longer a chat runs, the worse it gets

A single ChatGPT conversation degrades as it lengthens, and the fix is a fresh chat, not a longer one. Microsoft and Salesforce researchers ran more than 200,000 conversations through fifteen models in 2025 and found performance dropped an average of 39% once a task was teased out across a back-and-forth instead of asked in one go (the study). Worse, once it takes a wrong turn it tends to stay lost: it latches onto an early wrong assumption and every reply after that builds on it.

You feel this as a chat that was sharp for twenty minutes and is now stubbornly missing the point no matter how you rephrase. Correcting it in the same thread rarely pulls it back. Opening a new chat and giving it one clean, complete instruction usually does, because you’ve dropped the derailed history it was dragging along.

It doesn’t know your business

ChatGPT only knows the internet’s average answer to a question like yours, not the specifics of your business. It hasn’t seen your customers, your margins, your suppliers, or the three things that went wrong last time someone tried what it’s now confidently recommending. So it gives you the generic playbook in a self-assured voice, which is genuinely useful as a first draft and a sounding board, and misleading as a final decision.

The part it can’t do is the judgement that hangs on your particulars, and that’s exactly the part it will cheerfully pretend to have. Use it to widen your options and pressure-test your thinking. Keep the call that depends on knowing your own numbers where it belongs, with you.

Lean on it too hard and your own edge dulls

Outsource the thinking to ChatGPT and your own skill fades, which is the con that costs the most over time and the one people notice last. A small and much-argued-over 2025 MIT Media Lab preprint wired up 54 people writing essays, some using ChatGPT, some a search engine, some nothing, and tracked their brains on EEG. The ChatGPT group showed the lowest brain engagement, felt the least ownership of what they’d made, and couldn’t reliably quote the essay they’d finished minutes earlier; across four months they “consistently underperformed at neural, linguistic, and behavioral levels” (Your Brain on ChatGPT). It hasn’t been peer-reviewed and the authors themselves warned against over-reading it, but the direction matches what anyone who’s leaned hard on the tool already feels. They called it cognitive debt: the work feels done, but nothing got learned, so the capability isn’t there next time you need it without the tool.

The plainest version of this is a common refrain in Reddit’s cons-of-ChatGPT threads: lean on it for a task and you stop building the skill you’re handing over. For an owner the line to hold is simple. Let it draft, and let it argue with you. Keep doing the thinking you’d need to defend out loud in a room, because that’s the muscle you can’t afford to let waste.

The polished dud that lands on your team

There’s a second-order version of the confident-wrong problem, and it shows up the moment more than one person on a team is using AI. The name for it, coined in a September 2025 Harvard Business Review piece by researchers at BetterUp and Stanford, is workslop: output that looks finished but is hollow, and quietly shoves the real work onto whoever receives it. They put numbers on it. In a survey of 1,150 US workers, 41% had been handed a piece of it in the past month, and each one took an average of one hour and 56 minutes to untangle, which they costed at roughly $186 per person a month. The bit that should worry an owner isn’t the time, it’s the trust: 42% of people who received workslop rated the sender as less capable, and a third didn’t want to work with them again. AI-polished work you haven’t actually checked or thought through doesn’t read as productive to the person downstream. It reads as handing them your job.

Three more limits worth a line

Three further cons are real, and each has a guide of its own, so here they are in a sentence.

It agrees with you too readily, and will talk you into a shaky plan if you let it; the ways of framing a question so it pushes back instead are in how to use AI in your business. On a personal account it trains on your chats by default and a court has already reached ones people had deleted, which is the whole of is ChatGPT safe for your business. And its default writing is fluent and forgettable, samey enough to spot across a room, which is why for anything that has to sound like you it’s the weakest of the big tools and one of the alternatives worth switching to does that job better.

The verdict

None of this is a reason to drop ChatGPT. It’s a map of where to keep your hand on the wheel. The cons gather in one spot: the moment it produces a fact, a figure, a date or a decision, and does it in a voice that sounds equally sure whether it’s right or invented. So use it as the fastest first-drafter and thinking partner you’ve had, make it search for anything that changes, run its maths through a tool, and check anything with a name, a number or a consequence before you act. Match the job to the tool, keep the judgement that needs your specifics in your own hands, and it stays the best general assistant a small business can reach for. For the occasional job where it’s simply the wrong tool, the rest of the shortlist covers what to pick up instead.

Questions people ask

What are the main disadvantages of ChatGPT for a business?
The costly ones all sit around trust, not features. It invents facts and states them as confidently as real ones, its knowledge is months out of date, it can't reliably do maths on its own, it drifts the longer a chat runs, and it doesn't know anything specific about your business. Lean on it too hard and it quietly erodes your own skill. None of these stop you using it; they tell you where to check its work.
Can I trust what ChatGPT tells me?
Not blindly. Treat it like a fast, fluent, sometimes-wrong assistant. Anything you'd act on that carries a fact, a figure, a citation or a date needs checking, because the tool sounds exactly as sure when it's inventing as when it's right. Use it for drafts, structure and thinking; verify anything with a consequence attached.
Why does ChatGPT make things up, and is it getting worse?
It makes things up because it predicts plausible text rather than checking truth, and OpenAI's own research says training rewards a confident guess over admitting it doesn't know. It isn't a straight line: OpenAI's o3 reasoning model hallucinated more than the model before it, then GPT-5 cut factual errors sharply again. But the tendency is baked into how the tool works, so no release has taken it to zero, and it never flags when it's guessing.
Should I stop using ChatGPT because of these limits?
No. It's still the most useful general assistant a small business can reach for. The limits are a map of where to keep your hand on the wheel: check its facts and figures, make it search for anything current, and keep the judgement that needs your specifics in your own hands. For the odd job where it's the wrong tool, another model usually does that one thing better.

Rather have it built for you?