Teaching a data agent to say I don't know, and why that is hard
On a benchmark that prices wrong SQL, abstaining from everything scores 50% and the best real systems score 29.8% to 54.5%.
6 articles
On a benchmark that prices wrong SQL, abstaining from everything scores 50% and the best real systems score 29.8% to 54.5%.
A Google Ads campaign name takes 256 characters with no documented content validation. Meta documents no maximum at all.
Row-level security fails at the role the agent connects with. Per-tenant isolation moves the boundary out of application code and into architecture.
An AI agent can scan your whole warehouse without breaking a rule. Adding a LIMIT does not reduce what you are billed on non-clustered tables.
A semantic layer is where your metric definitions live in a form a machine reads before it computes. Not a tool: a set of written decisions.
Metric hallucination is AI computing correctly on a definition nobody wrote down. Supplying the definition raised accuracy from 45.5% to 67.7%.