The framing is a doctor watching a patient after surgery. The patient looks fine. The monitor is tracking heart rate, oxygen, and blood pressure so that something going wrong gets caught before it becomes a crisis. AI agents fail the same quiet way. Traditional software crashes, and a crash is loud. A probabilistic system keeps answering, and the same question asked twice can come back two very different ways, so quality can slide for weeks without anything looking broken. One minute on what to watch instead.
What this video covers
- Why a probabilistic system fails quietly rather than crashing, and what that hides from you
- The signals observability reads, whether answers are grounded in the right sources, whether quality is holding up, whether responses are slowing down, and whether token use and tool calls are running higher than expected
- Why those signals only mean something read together, as a view of whether the agent is behaving the way it was designed to
Chapters
- 0:00 The patient after surgery
- 0:22 Why AI does not fail the way software fails
- 0:44 The signals worth watching