Back to PagerDuty
PagerDuty logo
Sentry logo
PagerDuty + Sentry · LatchLoop

AI agent workflow: Add Sentry evidence to a PagerDuty incident

Give responders an incident timeline plus application-level evidence.

Workflow outcome

Give responders an incident timeline plus application-level evidence.

How an AI agent can give responders an incident timeline plus application-level evidence

This workflow gives an AI agent a defined job, a bounded set of records, and a result a person can review. The agent reads the relevant PagerDuty context and matches it with Sentry, applies the rules in the prompt, and keeps the source behind every recommendation. It returns a proposed handoff rather than taking consequential actions on its own.

Can an AI agent give responders an incident timeline plus application-level evidence?

Yes. Start with the scope, date range, decision rules, and fields that identify the right records. The agent can collect the evidence, compare states or sources, mark conflicts and missing data, and organize the result around the outcome above. A reviewer then checks the matches and judgment calls before approving messages, record updates, bookings, purchases, publishing, or other write actions. The guide below shows the records, boundaries, prompt, and handoff needed for this specific workflow.

Correlate time and service identity

PagerDuty supplies incident timing, affected service, alerts, responders, and response state. Sentry supplies error groups, events, traces, releases, and environment context. Used together, they can test whether an application error explains the alert and give responders a path from page to failing request.

Match on service, environment, and a tight time window. Request IDs, trace IDs, release versions, or error fingerprints provide stronger evidence than timestamps alone. The agent should redact personal data, authorization values, and request bodies that do not belong in an incident report.

Example starter prompt

Compare PagerDuty incident [ID] with Sentry evidence for [project and environment] between [start] and [end].

Build the PagerDuty timeline, then identify Sentry error groups, events, traces, and releases that overlap the affected service and symptom. Cite incident events and Sentry identifiers. Label correlations and confirmed causal evidence separately.

Do not change incident state, assign issues, or deploy code. Return the leading evidence, counterevidence, and next check.

Test the leading explanation

Look for affected and unaffected requests around the same time. A release that appears shortly before the page is a lead; a matching failing trace tied to the user symptom is stronger. Preserve alternative causes until the evidence rules them out.

The handoff should include safe links to PagerDuty and Sentry, the first confirmed divergence, owner, and unresolved data gaps.

Questions this workflow answers

Which application errors and traces best explain the user impact behind an on-call alert?

PagerDuty supplies the incident, alerts, service, responders, status, and timeline. Sentry supplies exceptions, releases, traces, affected requests, and environment context. The agent aligns them by service, environment, incident window, release, and user-visible symptom before suggesting a relationship.

Start with the first confirmed impact and compare affected and unaffected requests around that time. The agent can cluster relevant Sentry events, identify when the pattern began, and show whether the error volume or trace behavior changes with the PagerDuty alert. A release just before the page is a lead; a matching failing trace on the affected path is stronger evidence.

Ask for support and contradiction for each hypothesis. The alert may result from a downstream service, infrastructure condition, or monitoring rule rather than the most visible exception. Preserve issue, trace, release, incident, and alert links while redacting tokens and personal data. Missing observability belongs in the gap list.

The handoff includes a shared timeline, first divergence, candidate error pattern, impact, owner, next checks, and unresolved data gaps. Responders decide whether to change incident status, deploy, or communicate. The agent does not resolve the page or claim causation from timing alone. It narrows the evidence so on-call effort starts with the requests most closely tied to the observed failure.

Sampling should cover the alert’s actual population. If the page reports checkout failures, the agent can compare failing and successful traces for the same route, environment, release, and period. It should show whether one exception dominates, whether it begins before or after the alert threshold, and whether affected requests share a tenant, region, browser, or dependency. A loud background error that occurs equally in successful requests is weak evidence. That contrast helps responders choose the trace that can explain impact, not merely the issue with the largest event count.

Get Started

Build as fast as you can think.

LatchLoop works where you do to build with you.