Debugging Agents in Different Environments - Live Workshop (Sponsored)Your agent returns something odd. Was it the prompt, a tool call that timed out, or a response your code could not parse? Without traces, you are guessing. In this hands-on workshop, Serge from Sentry instruments three agents with Sentry Agent Tracing: a chatbot in an ecommerce store, a custom Slack agent, and a GitHub Action that reviews PRs. You’ll see how to catch bad tool calls and unexpected output, plus how to track token spend and performance across every agent you run. A customer asks a company’s new AI-based support assistant whether a subscription purchased 15 days ago qualifies for a refund. Let’s assume that the assistant responds immediately by explaining that the company offers a 30-day refund window, describes the cancellation process, and promises that the money will arrive within five working days. All of it sounds quite helpful, confident, and complete. But there is just one problem. In reality, the company only allows refunds within 14 days. Also, there is no promise about the 5-day processing time. The assistant has actually taken bits of a plausible policy and invented a fake commitment out of thin air. In other words, it has lied. The customer now expects something that the company never offered. Though this example might sound fictitious, it shows one of the most important problems in applications built with LLMs. While LLMs can produce excellent language, they can get the underlying information totally wrong. In this article, we will look at why this problem of hallucinations happens with LLMs and the techniques that can help make LLMs more dependable for answering. Here’s what we will cover:
What Hallucination Actually MeansHallucination in the context of LLMs is generated information that is factually incorrect, invented, or inconsistent with the material the model is supposed to use. Here’s how a hallucination is happens: The entire response may not be wrong. In fact, an otherwise useful explanation can contain a single fabricated date, an unsupported promise, or a reference to a document that doesn’t exist. That small detail may be the part the reader relies upon. We can’t outright call this behavior lying, but it comes pretty close to a lie. The answer resembles a confident falsehood. However, lying usually implies an intention to deceive. A hallucination doesn’t establish that intention. The model can produce an incorrect answer through its normal generation process. In the support example from earlier, the immediate problem is that the application presents an incorrect generated policy as established company information. The context matters a lot here. Inventing a refund policy for a fictional company is appropriate when the task asks for one. However, presenting that same invention as the policy of an actual business is a factual failure. Ultimately, the difference is more about what the answer claims to represent. Three Ways an Answer Can Go WrongThe AI support assistant’s errors become easier to reason about when divided into three useful categories:
These three categories overlap. |