AI Agents in Production: Silent Failures, Ghost Runs, and How to Catch Both
AI agents fail differently from regular software. They don't crash — they drift, loop, stall, and consume resources while producing nothing. This guide covers the two failure modes unique to agentic AI and how to monitor for both.
Your AI agent is running. The task was assigned 3 hours ago. The process is active. Tokens are being consumed.
No output has been produced.
Welcome to the ghost run — the silent failure mode unique to agentic AI that your conventional monitoring stack has no concept of.
Regular software fails loudly: exceptions, stack traces, 500 errors, process crashes. You know something broke. You fix it.
AI agents fail quietly. They don't crash — they drift. They loop. They stall on a tool call. They produce structurally valid output that's semantically wrong. They consume your entire token budget on a task that should have taken 400 tokens and cost you $0.08. They run for hours, look busy, and deliver nothing.
This guide covers both failure modes — silent failures and ghost runs — what causes them, why they're invisible to conventional monitoring, and how to catch them before they cost you.
Two Failure Modes Unique to Agentic AI
Silent Failures
A silent failure in an AI agent is any failure that produces no visible error signal. The agent process is running. No exception was thrown. No alert fired. But the task didn't complete correctly — or didn't complete at all.
Silent failures in agents include:
- Tool call fails, agent retries silently until it hits a loop or token limit
- External API times out, agent stalls waiting for a response that never comes
- Agent produces output that passes validation but is factually wrong or structurally incomplete
- Human-in-the-loop pause never gets resolved — agent waits indefinitely
- Memory or context window fills, agent silently truncates earlier context and makes decisions on incomplete information
Ghost Runs
A ghost run is a specific type of silent failure where an agent is actively executing — consuming compute, tokens, and API budget — but producing no useful output.
The name captures the experience: the process is there, it looks alive, but there's nothing behind it. Your monitoring shows activity. Your OpenAI dashboard shows token consumption. Your agent is technically "running." But the task it was supposed to complete is never going to be delivered.
Ghost runs happen when:
- An agent enters an infinite loop — calling the same tool repeatedly, getting the same result, trying again
- An agent is stuck waiting for a resource that will never become available
- An agent's reasoning has gone off-track and it's exploring an irrelevant path indefinitely
- An agent hit its context limit and is making increasingly incoherent decisions with no awareness that it's lost the thread
- A chained multi-agent workflow produced a bad intermediate output and the downstream agent is hallucinating on corrupted input
The difference between a silent failure and a ghost run: a silent failure means the agent stopped. A ghost run means the agent kept going — in the wrong direction, expensively, producing nothing.
Why These Failures Are Invisible to Conventional Monitoring
Understanding why your existing tools miss this is important — because the instinct is to add more logging, and logging alone doesn't solve it.
Uptime monitors check if your server responds. Your agent process is running. ✅ Green.
Error monitors watch for exceptions and 5xx responses. A looping agent doesn't throw exceptions — it makes valid API calls in a loop. ✅ Green.
Log monitors watch for error-level entries. An agent stalling on a tool call logs the tool call, not an error. ✅ Green.
Token usage dashboards show consumption over time. A ghost run shows high consumption. You might notice — but you don't know which agent, which task, or why. And you're looking at historical data, not a real-time alert.
LLM observability tools (Langfuse, Helicone, Arize) trace individual LLM calls — prompts, completions, latency. Excellent for debugging after the fact. But they don't alert you in real time when an agent has been running for 3 hours without completing, or when token consumption for a single task is 10x your baseline.
The gap: all these tools watch what the agent is doing. None of them watch whether the agent is getting anywhere.
That's the monitoring question agentic AI requires: "Is this agent making progress toward completing its task?"
The Ghost Run Taxonomy
Not all ghost runs are the same. Understanding the type helps you instrument the right detection:
Type 1 — The Infinite Loop
The agent calls a tool, gets a result, decides to call the same tool again with slightly different parameters, gets a similar result, calls it again. This continues until max_iterations is hit or the token budget is exhausted.
Signature: high tool call count, low unique tool diversity, near-zero progress per iteration, task never completed.
Detection: track tool call frequency per unique tool. Alert if the same tool is called more than N times in a single task run without a completion event.
Type 2 — The Stall
The agent is waiting for an external resource — an API response, a human approval, a file to be available — that never comes. The agent process is alive. It's waiting. Nothing is happening.
Signature: no activity for extended period, task not completed, no error logged.
Detection: track time since last meaningful event (tool call, LLM call, state update). Alert if no activity for longer than your baseline completion time.
Type 3 — The Rabbit Hole
The agent's reasoning goes off-track. It starts pursuing a tangential path — researching something adjacent to the task, generating outputs that don't address the original objective, calling tools that have nothing to do with the assigned work.
Signature: high activity, high token consumption, zero relevant output, task never completed.
Detection: hardest to catch automatically. Requires output validation — comparing the agent's output to the original task objective — or human review of completion events.
Type 4 — The Context Collapse
The agent hits its context window limit. Earlier context — including the original task instruction — is truncated. The agent continues operating, but it's making decisions without remembering what it was supposed to do. It may complete something — just not the right thing.
Signature: task appears to complete (completion event fires), output is structurally valid, output doesn't address the original objective.
Detection: track token count throughout the run. Alert when context window utilisation exceeds 80%. Require human review of completion events on long-running tasks.
Type 5 — The Cascade
In multi-agent workflows, one agent produces bad output and passes it downstream. The downstream agent receives corrupted input and tries to work with it — hallucinating, producing garbage, or stalling. Each subsequent agent in the chain makes the problem worse.
Signature: downstream agents completing faster than normal (short-circuiting on bad input) or much slower (stalling on corrupted context), final output wrong or missing.
Detection: validate output at each handoff point between agents. Track inter-agent completion times. Alert on statistical deviation from baseline handoff latency.
Instrumenting AI Agents for Silent Failure Detection
The key insight: you need to track not just what the agent is doing but whether it's making progress. Progress is measured in completion events — tool calls that advance the task, milestones hit, output produced.
Install NotiLens
pip install notilens langchain langchain-openai
import notilens
nl = notilens.init(
name="ai-agents",
token="YOUR_TOKEN",
secret="YOUR_SECRET"
)
The Core Instrumentation Pattern
Every agent task needs four instrumented points:
# 1. Task assigned — flow starts
run = nl.task("agent-task")
run.start()
# 2. Progress checkpoints — agent is making progress
run.progress("Tool: web_search — querying for X")
run.metric("tool_calls", 1) # accumulates
run.metric("tokens", tokens_used) # accumulates
# 3. Task completed — flow ends successfully
run.output_generated(f"Output: {result[:120]}")
run.complete("Task completed successfully")
# 4. Task failed — flow ends with failure
run.fail(f"Agent failed: {str(error)}")
If run.start() fires but run.complete() never comes — broken flow detection catches it. If run.start() never fires — Smart Silence Detection catches it. If both fire but token consumption or tool call count is anomalous — ML anomaly detection catches it.
Full LangChain Agent Instrumentation
import notilens
import time
from langchain.callbacks.base import BaseCallbackHandler
from langchain.schema import AgentAction, AgentFinish, LLMResult
from typing import Any, Dict, List, Union
nl = notilens.init(name="ai-agents")
class NotiLensAgentHandler(BaseCallbackHandler):
"""Production monitoring handler for LangChain agents."""
def __init__(self, task_name: str = "agent"):
self.task_name = task_name
self.run = None
self.iteration_count = 0
self.token_count = 0
self.cost_usd = 0.0
self._tool_history = [] # for loop detection
self._start_time = None
# ── Lifecycle ─────────────────────────────────────────────────────────────
def on_chain_start(self, serialized: Dict, inputs: Dict, **kwargs):
self.run = nl.task(self.task_name)
self._start_time = time.time()
self.run.start()
self.run.track("agent.started", "Agent chain started",
meta={"input": str(inputs)[:200]})
def on_agent_finish(self, finish: AgentFinish, **kwargs):
output = str(finish.return_values.get("output", ""))
elapsed = round(time.time() - self._start_time, 2)
self.run.metric("runtime_seconds", elapsed)
self.run.output_generated(output[:200])
self.run.complete("Agent task completed")
def on_chain_error(self, error: Union[Exception, KeyboardInterrupt], **kwargs):
self.run.fail(f"Chain error: {str(error)}")
# ── Tool calls ────────────────────────────────────────────────────────────
def on_agent_action(self, action: AgentAction, **kwargs):
self.iteration_count += 1
self._tool_history.append(action.tool)
self._tool_history = self._tool_history[-5:] # keep last 5
self.run.metric("tool_calls", 1)
self.run.progress(
f"[{self.iteration_count}] Tool: {action.tool}",
meta={"input": str(action.tool_input)[:120]}
)
# ✦ Ghost run detection — loop detection
# Same tool called 3+ times in last 5 calls = likely stuck
if len(self._tool_history) >= 3:
recent = self._tool_history[-3:]
if len(set(recent)) == 1:
self.run.loop(
f"Ghost run detected — '{action.tool}' called "
f"{self.iteration_count} times without progress"
)
def on_tool_start(self, serialized: Dict, input_str: str, **kwargs):
tool_name = serialized.get("name", "unknown")
self.run.track("tool.started", f"{tool_name}: {input_str[:80]}")
def on_tool_end(self, output: str, **kwargs):
self.run.track("tool.completed", str(output)[:120])
def on_tool_error(self, error: Union[Exception, KeyboardInterrupt], **kwargs):
self.run.error(f"Tool error: {str(error)}")
# ── LLM calls ─────────────────────────────────────────────────────────────
def on_llm_end(self, response: LLMResult, **kwargs):
usage = (response.llm_output or {}).get("token_usage", {})
tokens = usage.get("total_tokens", 0)
if tokens:
self.token_count += tokens
cost = tokens * 0.0000002 # gpt-4o-mini rate
self.cost_usd += cost
self.run.metric("tokens", tokens)
self.run.metric("cost_usd", round(cost, 6))
# ✦ Ghost run detection — token budget alert
if self.token_count > 50_000:
self.run.error(
f"Token budget warning — {self.token_count:,} tokens consumed "
f"(${self.cost_usd:.4f}) without task completion"
)
def on_llm_error(self, error: Union[Exception, KeyboardInterrupt], **kwargs):
self.run.error(f"LLM error: {str(error)}")
# ── Usage ─────────────────────────────────────────────────────────────────────
from langchain_openai import ChatOpenAI
from langchain.agents import AgentExecutor, create_openai_tools_agent
handler = NotiLensAgentHandler(task_name="research-agent")
llm = ChatOpenAI(model="gpt-4o-mini", callbacks=[handler])
agent = create_openai_tools_agent(llm, tools, prompt)
agent_executor = AgentExecutor(
agent=agent,
tools=tools,
callbacks=[handler],
max_iterations=15
)
try:
result = agent_executor.invoke({"input": user_input})
except Exception as e:
if handler.run:
handler.run.fail(str(e))
raise
Detecting Stalls With Smart Silence Detection
For agents assigned tasks that should complete within a predictable window, Smart Silence Detection handles stall detection automatically.
Each run.progress() call resets the silence timer. If no progress event arrives for longer than your agent's learned baseline completion time — NotiLens alerts.
You don't configure a manual timeout. The ML model learns how long your agent typically takes to produce a tool.completed or agent.finished event, and alerts when the gap is genuinely anomalous.
For agents with strict SLAs, you can also set a manual silence window:
# Example: agent must complete within 30 minutes
# Create topic with manual silence window = 35 minutes
# run.progress() calls keep the window alive
# If no progress for 35 min — alert fires
run.progress("Starting tool search phase") # resets timer
# ... agent runs ...
run.progress("Tool results received, synthesising") # resets timer
# ... agent completes or stalls ...
run.complete("Done") # resolves the flow
Human-in-the-Loop Stall Detection
If your agent pauses for human input, instrument the pause explicitly:
async def run_agent_with_review(task: str):
run = nl.task("hitl-agent")
run.start()
try:
# Phase 1: agent generates draft
run.progress("Generating draft output")
draft = await generate_draft(task)
run.output_generated(f"Draft: {draft[:120]}")
# ✦ Human approval required — urgent alert fires
run.input_required(
"Human review required — agent paused awaiting approval",
meta={"draft_preview": draft[:200]}
)
# Wait for approval (your mechanism — webhook, UI, CLI)
approved, reviewer = await wait_for_approval(draft)
if approved:
run.input_approved(f"Approved by {reviewer}")
result = await execute_approved_draft(draft)
run.complete("Task completed after human approval")
else:
run.input_rejected(f"Rejected by {reviewer}")
run.cancel("Task cancelled — human rejected draft")
except asyncio.TimeoutError:
# Nobody approved within SLA
run.timeout("Human approval timeout — no response within SLA")
except Exception as e:
run.fail(str(e))
run.input_required() sends an urgent push notification to your phone with 5-minute repeat reminders until acknowledged. A human-in-the-loop pause that nobody responds to never goes unnoticed.
Multi-Agent Workflow Monitoring
For agent pipelines where one agent hands off to another, track each agent as its own task and validate the handoff:
import notilens
import json
nl = notilens.init(name="agent-pipeline")
async def run_pipeline(user_request: str):
pipeline_run = nl.task("pipeline")
pipeline_run.start()
try:
# ── Agent 1: Research ──────────────────────────────────────────────
research_run = nl.task("research-agent")
research_run.start()
research_output = await research_agent.run(user_request)
# Validate handoff output before passing downstream
if not research_output or len(research_output) < 50:
research_run.fail("Research output too short — possible ghost run")
pipeline_run.fail("Pipeline aborted — bad research output")
return
research_run.metric("output_length", len(research_output))
research_run.complete("Research phase complete")
# ── Agent 2: Synthesis ─────────────────────────────────────────────
synthesis_run = nl.task("synthesis-agent")
synthesis_run.start()
synthesis_output = await synthesis_agent.run(research_output)
if not synthesis_output:
synthesis_run.fail("Synthesis produced no output — cascade failure")
pipeline_run.fail("Pipeline aborted — synthesis failed")
return
synthesis_run.metric("output_length", len(synthesis_output))
synthesis_run.complete("Synthesis phase complete")
# ── Pipeline complete ──────────────────────────────────────────────
pipeline_run.output_generated(synthesis_output[:200])
pipeline_run.complete("Pipeline completed successfully")
return synthesis_output
except Exception as e:
pipeline_run.fail(f"Pipeline error: {str(e)}")
raise
Each agent in the pipeline gets its own run context, its own token metrics, and its own broken flow detection. A cascade failure — where Agent 1 produces bad output that breaks Agent 2 — is visible immediately rather than only detectable in the final output.
The AI Agent Monitoring Checklist
Before shipping an agent to production:
-
run.start()fires when task is assigned -
run.progress()fires on every meaningful step (tool call, milestone) -
run.complete()fires on successful task completion with output -
run.fail()fires on any unhandled exception -
run.error()fires on tool errors (non-terminal) -
run.loop()fires when repeated tool calls detected -
run.timeout()fires when agent stalls beyond expected window -
run.input_required()fires when human approval needed -
run.metric("tokens", ...)tracks token usage per run -
run.metric("cost_usd", ...)tracks cost per run -
run.metric("tool_calls", ...)tracks iteration count - Smart Silence Detection active — alerts if task assigned but no progress
- Broken flow detection active — alerts if task starts but never completes
- Token budget alert configured — fires if token count exceeds threshold
- On-call routing configured for ghost run alerts
- Tested — deliberately triggered a loop and confirmed alert fired
- Tested — deliberately stalled an agent and confirmed silence alert fired
For how AI agent monitoring fits into a complete founder stack alongside payments, cron jobs, and infrastructure, see The Founder's Monitoring Stack.
What You See in NotiLens for a Ghost Run
A healthy agent run:
✅ agent.started Research agent — task assigned
⚙️ tool.started web_search: latest AI regulation news
✅ tool.completed Tool output: [12 results found]
⚙️ tool.started read_url: https://...
✅ tool.completed Tool output: [article content]
✅ output.generated Summary: EU AI Act passed in March...
✅ task.completed Agent task completed
tokens: 3,240 | cost: $0.0006 | tool_calls: 4 | runtime: 18s
A ghost run — loop detected:
✅ agent.started Research agent — task assigned
⚙️ tool.started web_search: AI regulation
✅ tool.completed Tool output: [12 results]
⚙️ tool.started web_search: AI regulation news ← same tool
✅ tool.completed Tool output: [11 results]
⚙️ tool.started web_search: latest AI regulation ← same tool
🔄 ghost_run.detected 'web_search' called 3 times without progress
→ Push notification fired → On-call paged
tokens: 12,400 | cost: $0.0025 | tool_calls: 8 | runtime: 4m 12s
A stall — silence detection:
✅ agent.started Research agent — task assigned
⚙️ tool.started database_query: SELECT...
[no further events for 47 minutes]
🔇 smart_silence.fired No progress in 47 min — agent stalled
→ Push notification fired → Escalation policy triggered
Summary
AI agents fail differently from regular software — and conventional monitoring was built for regular software.
Silent failures happen when agents stall, loop, or produce wrong output with no visible error signal. Ghost runs happen when agents keep running — consuming compute and budget — while producing nothing useful.
Catching both requires monitoring that watches for progress, not just activity. The NotiLens agent instrumentation pattern — run.start(), run.progress(), run.complete(), run.fail() — gives your agent a heartbeat that conventional monitoring tools can't provide.
One caught ghost run on a complex research agent pays for months of NotiLens.
Try AI agent monitoring free for 7 days — no credit card required.
Frequently Asked Questions
What's the difference between a ghost run and a hallucination? A hallucination is when an LLM produces factually incorrect output — it completes the task, but the content is wrong. A ghost run is when the agent never produces meaningful output at all — it's running but getting nowhere. Hallucinations are an output quality problem. Ghost runs are an infrastructure and monitoring problem. Both are costly but require different solutions.
Does NotiLens work with agent frameworks other than LangChain?
Yes. The NotiLens SDK is framework-agnostic. The same run.start(), run.progress(), run.complete(), run.fail() pattern works with CrewAI, AutoGen, Pydantic AI, LlamaIndex agents, and any custom agent loop. The callback handler approach is LangChain-specific, but the core SDK works anywhere.
How do I set the right token budget threshold for ghost run detection? Start with 3–5x your agent's average token consumption per successful run. If your agent normally uses 3,000–5,000 tokens, set the ghost run alert at 15,000–20,000. Adjust after a week of data. Smart Silence Detection will also automatically learn your normal token consumption range and alert on statistical anomalies — so you don't have to guess the right threshold.
Can I monitor agents that run for hours or days?
Yes. Long-running agents benefit most from run.progress() calls at regular checkpoints — every completed subtask, every milestone, every tool phase. Smart Silence Detection learns the normal gap between progress events and alerts if the gap becomes anomalously large — catching stalls mid-run without requiring you to define a fixed timeout.
What's the performance impact of the NotiLens callback handler? Negligible. The NotiLens SDK calls are non-blocking and fire asynchronously. In a typical agent run, the NotiLens overhead is measured in milliseconds compared to LLM call latency measured in seconds. The monitoring cost is effectively zero relative to the agent's actual compute cost.
How do I monitor agents that run in parallel across many workers?
Each parallel agent run creates its own nl.task() run context. NotiLens tracks each independently. If you're running 50 agents in parallel, you get 50 independent run timelines. You can filter by task name in the NotiLens dashboard to see all concurrent runs of a specific agent type, or look at the aggregate anomaly detection to catch cases where an unusually high percentage of parallel runs are failing.