Vinit Agrawal, co-founder and CTO at Tars, builds this around one line on an enterprise RFP: how do you verify the accuracy of responses by your AI agents? For fifty years that question had a boring answer. Software was rules, rules could be tested, and a green test suite meant the thing worked. An LLM based agent does not work that way. It predicts rather than follows instructions, and that is a property of anything that speaks a human language, not a flaw waiting to be patched in a later model.
The argument runs on an invented case, two companies he calls A and B, both shipping the same kind of e-commerce support agent, both hit by the same incident. One treats it as a bug. The other builds a way to know. Watch this if you have to answer for an agent's accuracy in front of a client, a regulator, or a board.
What this video covers
- Why an agent writes every answer on behalf of someone else, and what that puts at risk for the brand it speaks for
- The story starts with a refund promised to a shopper who was not eligible, with no line of code to point at and no bug to fix
- Why test suites stopped answering the accuracy question once the software started predicting instead of following rules
- What happens to the company that treats it as a bug, edits the prompt, and stops looking
- How the other company reads and hand-labels real conversations first, then automates the reading with a second model that scores every conversation
- Why the hand-labeled set is what stops the who-checks-the-checker problem from running forever
Chapters
- 0:00 The question every AI company gets asked
- 0:35 Two companies, one product, answering in someone else's name
- 1:46 How AI products fail, quietly and at scale
- 2:31 Why software testing stopped working
- 3:35 Guessing versus measuring
- 4:42 Automating the looking, and who judges the judge
- 5:50 What a number you can trust changes
- 7:24 The new moat