Every customer-facing agent runs a loop. Call the model with the user's message, decide which tool to use and with what input, run it, read the output, decide the next step, and keep going until the request is fulfilled. The tool that decides answer quality most often is knowledge retrieval, so evaluating an agent starts with evaluating what it retrieves. That is the subject here.
Vinit Agrawal defines good retrieval on three axes, shows why only two of them are really variables, and then fixes the thing that usually breaks a retrieval test set. Label your ground truth as a list of chunk IDs and it dies the moment you change chunking strategy. Anchor it to character positions in the document instead and one golden dataset survives any retriever you point at it. Watch this if you are choosing between retrieval setups and want a measurement instead of an impression, and particularly if you work somewhere an incomplete answer is a compliance problem rather than just a bad experience.
What this video covers
- The agent loop in plain terms, and why knowledge retrieval is the tool that decides answer quality most often
- Why retrieval needs its own evaluation system once you have more than one retriever configured on the same Knowledge Base
- Good retrieval defined on three axes, accuracy, relevance and completeness, and why accuracy is not really the variable
- How relevance and completeness formalize into precision and recall by comparing retrieved chunks against the snippets that hold the answer
- Why ground truth stored as chunk IDs becomes useless the moment you change chunking, and why that means rebuilding the dataset for every setup you test
- How anchoring ground truth to character spans, a start index and an end index in the document itself, lets one dataset score any retriever that reports where its chunks begin and end
Chapters
- 0:00 The agent loop and where retrieval sits in it
- 1:27 Why retrieval needs an evaluation system of its own
- 2:54 Accuracy, relevance, and completeness
- 3:39 Recall and precision explained
- 4:45 Why ground truth built on chunk IDs breaks
- 5:53 Anchoring ground truth to character spans
- 7:23 Recap of the three ideas