An assistant that rounds a balance or invents a rate isn’t a minor error: in regulated banking, every word can end in a complaint, a fine, or a lawsuit.
Ariel messages his bank’s chatbot to ask how much he’d be charged for paying off his loan early. The reply arrives in seconds, sounding certain: “the penalty is 1% of the remaining balance.” Ariel does the math, decides it’s worth canceling, and transfers the money. A month later, his statement shows a different charge: 4%, not 1%. He files a complaint, and it takes the bank weeks to explain where the number the assistant gave him came from. No one knows for sure, because no one designed that number: the model generated it because it sounded like a reasonable answer.
When the model fills in what it doesn’t know
Language models don’t fail the way a traditional program fails, with a visible error that halts execution. They fail by filling in: when they don’t have the exact data, they generate the most probable data based on the pattern of the text, with the same confidence they’d have if it were correct. For a movie recommendation, that habit is almost a virtue. For a rate, a due date, or a balance, it’s a problem, because the customer has no way to tell a verified answer from an invented one: both sound equally certain.
In banking, there’s no such thing as a small error
That same behavior, tolerable in other industries, changes category in banking. Misstating a rate, a term, or a condition isn’t an anecdote: it’s regulated information, with consumer protection bodies and financial regulators that exist precisely for this. A customer who made a decision based on false information has a legitimate claim, and the bank answers for what its assistant said with the same seriousness it would answer for what an employee said at a branch. The difference is that an employee can recall why they said what they said. A model that generated a figure out of nowhere cannot.
No one can reconstruct where the answer came from
That’s where the second, quieter problem shows up: when the complaint comes in, someone at the bank has to be able to explain how that answer was reached. If the number came from an unconstrained generative process, there’s nothing to reconstruct: there was no query to a system, no rule that was applied, just a statistical inference over text. Audit and compliance are left with nothing to show the regulator beyond “the model got it wrong,” which isn’t an answer a bank can give.
The answer: every piece of data, backed by a source
With Singular, no agent invents a figure, because none generates one freely: every rate, balance, term, or condition that appears in a conversation is obtained by querying the bank’s systems live, not by completing a text pattern. For regulated topics, Studio lets the bank set in advance exactly what each agent can and can’t say, with official sources as the only possible origin for that answer: there’s no room for the model to improvise a number. Every conversation is logged end to end, so if a customer or a regulator asks why the assistant said what it said, the bank can show exactly which data and which rule generated it. And Comandante monitors every conversation live, so a questionable answer gets corrected before it reaches the customer, not after it’s already triggered a complaint.
The difference isn’t that Singular is more cautious: it’s that accuracy doesn’t depend on the model’s luck in completing the pattern correctly this time. Every answer can be verified, because every answer has real data behind it.
In the next article in the series: what happens when these conversations multiply by the thousands at the same time, and why scaling can’t mean losing quality?

