A plain list of what a language model cannot do, from not knowing your organisation to not being accountable, so staff can judge each task rather than rely on belief.
Most conversations about AI at work land on one of two extremes: it can do almost anything, or it is unreliable nonsense dressed up in confident prose. Neither extreme is useful when you are deciding whether to hand it a real task on a Tuesday afternoon. What is useful is a short, plain list of what it actually cannot do, so the decision is about the task in front of you rather than a belief about the technology in general.
None of what follows is a technical deep dive, and none of it is a reason to avoid the tools. It is the boundary line every member of staff needs, regardless of department.
A language model learns general patterns from a huge amount of text written by other people, at a point in the past. It has not read your safeguarding policy, your last board minutes, or the email thread where a decision was actually made, unless you put that text in front of it in the conversation. Ask it about "our policy" without pasting the policy and it will answer from what organisations like yours typically do, not from what yours actually says. That answer can be fluent, specific-sounding, and wrong.
Every model is trained up to a point in time, and its knowledge stops there unless it is explicitly connected to a tool that can search the live web. Ask about a regulation, a funder, or a news event from after that point and it may answer anyway, filling the gap with the closest pattern it has rather than saying "I don't know." The honest answer is rare unless you ask for it directly: tell me if you are not certain this is current.
We have written before about the jagged frontier: capability that is high in some directions and low in others, with no smooth boundary you can feel from inside the conversation. The tool's fluency does not vary with its accuracy. A confidently wrong answer reads exactly like a confidently right one, which is the actual reason to check outputs on unfamiliar tasks, not politeness or caution for its own sake.
It has read more than anyone in the building. It has met none of your beneficiaries.
Treat it as a fast, extremely well-read colleague who joined this morning, has never worked at your organisation, and cannot be asked to double-check their own homework. That framing answers most of the practical questions on its own: give it your actual documents rather than asking it to guess at them, ask it to flag uncertainty rather than assuming silence means confidence, and keep a named person accountable for anything that leaves the building.
Before using it for a task that touches a decision, ask one question: if this answer is confidently wrong, would anyone here notice? If the honest answer is no, that is the task to slow down on, not the task to trust most.
No external statistic cited; references the jagged-frontier framing established in tree #5 (already scheduled) rather than introducing new evidence.
If this is the question on your desk, a thirty-minute call tells you whether the service fits, or that you do not need us yet.