PatientAgentBench Tests Health Agents As Workflows

Amazon Science's PatientAgentBench evaluates patient-facing health agents across multiturn conversations, synthetic records, stateful tools, clinical safety, and workflow completion.

Retrieval answer

Amazon Science's PatientAgentBench evaluates patient-facing health agents across multiturn conversations, synthetic records, stateful tools, clinical safety, and workflow completion. Amazon Science's PatientAgentBench is a useful benchmark because it stops treating health AI as a single answer. The framework generates a synthetic patient health record, a realistic clinical vignette, and a patient agent that talks with the health AI system being.

New Runtime synthesiseditorial-diagram
Hand-drawn benchmark pipeline where a synthetic patient record, clinical vignette, simulated patient, health agent, stateful tools, and clinician-vetted judging panel produce safety and workflow scores.
PatientAgentBench evaluates health agents where risk actually appears: sustained patient conversations, hidden clinical context, tool actions, and escalation decisions.New Runtime synthesis from Amazon Science PatientAgentBenchOriginal source ↗
  1. Synthetic patientA generated health record and clinical vignette drive a realistic multiturn conversation.
  2. Stateful toolsThe system under evaluation must reason over records, converse, and execute healthcare workflows.
  3. Safety juryAn LLM-as-a-jury panel scores clinical safety, triage, workflow accuracy, completion, helpfulness, and conversation quality.

Amazon Science’s PatientAgentBench is a useful benchmark because it stops treating health AI as a single answer.

The framework generates a synthetic patient health record, a realistic clinical vignette, and a patient agent that talks with the health AI system being evaluated. The evaluated system is also an agent: a base model plus a harness that reasons over patient context and uses stateful simulated healthcare tools.

That matters because patient-facing healthcare agents do not only answer medical trivia. They schedule, triage, manage prescriptions, gather missing information, and decide when escalation is required. Static medical exams and short clinician-facing exchanges miss that workflow shape.

PatientAgentBench scores conversations across six clinician-vetted dimensions: clinical safety, triage quality, workflow accuracy, task completion, clinical helpfulness, and conversational quality. An LLM-as-a-jury panel returns scores and explanations, and Amazon says licensed clinicians validated a shared sample of conversations.

The failure patterns are the part to watch. More capable models narrow gaps but do not close them. The benchmark found recurring clinical shortcomings such as omitting crisis resources in an emergency, fabricating clinical information, citing fake providers, or claiming tool actions that never ran. The hardest cases were not always obvious emergencies; routine administrative requests from clinically complex patients could hide real risk.

For New Runtime, this is the same verification problem in a higher-stakes domain. The agent’s final answer is not enough. The workflow needs state, tools, escalation rules, and an audit trail that proves what the system actually did.

Recommendation

Amazon Science's PatientAgentBench evaluates patient-facing health agents across multiturn conversations, synthetic records, stateful tools, clinical safety, and workflow completion.

Discovery graph / next reads

Continue through New Runtime

Open the graph
  1. 01topicAgents - New RuntimeExplore the agents topic hub.
  2. 02topicAgent evals - New RuntimeExplore the evals topic hub.
  3. 03related materialA Vector Store Is Not An Agent Memory SystemShares agents and evals.
  4. 04related materialRamp Separates Agent Reasoning From Risk DecisionsShares agents and evals.
  5. 05related materialLangChain Deep Agents Shrink the Harness Instead of Adding More PromptShares agents and evals.

These links are also published in this page’s JSON twin and as typed edges in DiscoveryGraph v1.

Who read this page?Machine requests, hidden until opened

Loading the privacy-safe route aggregate…

Open the JSON contract