A confident answer with nothing behind it.
Making things up is not a bug that will be patched out. It is the same machinery that makes the model useful, pointed at a question it has no evidence for.

The thing is always guessing
A language model produces the next fragment that best fits what came before. That is the whole operation. When it tells you something true, it is because the true continuation was also the most likely one. The mechanism did not change when it was right.
So there is no internal switch between recalling and inventing. There is one process, running at the same confidence either way.
Why it sounds so sure
Fluency and accuracy are produced by different things, and only one of them is trained hard. The model is relentlessly optimised to sound like well-formed prose. Nothing in that optimisation rewards saying “I am not certain”.
This is why a fabricated citation reads exactly like a real one. The format is learned perfectly. The contents were never checked, because there was nothing to check them against.
The failure has a shape
It is most likely where the request is specific and the evidence is thin: exact dates, case numbers, page references, the name of a person who did a particular thing. Precisely the details a reader treats as proof the thing knows what it is talking about.
The fix is not a better model
You reduce this by changing the question from recall to lookup. Fetch the actual passages, hand them over, and require the answer to come from them. Then ask for the source alongside the claim, and check the source exists.
A model that cannot find supporting material should return nothing. An empty answer is a cheap problem. A confident wrong one is an expensive problem that arrives looking like a correct one.