{"schema_version":"newruntime-agent-readable-v0.2","type":"post","stable_id":"post:patientagentbench-health-agent-safety","slug":"patientagentbench-health-agent-safety","title":"PatientAgentBench Tests Health Agents As Workflows","description":"Amazon Science's PatientAgentBench evaluates patient-facing health agents across multiturn conversations, synthetic records, stateful tools, clinical safety, and workflow completion.","retrieval_nugget":"Amazon Science's PatientAgentBench evaluates patient-facing health agents across multiturn conversations, synthetic records, stateful tools, clinical safety, and workflow completion. Amazon Science's PatientAgentBench is a useful benchmark because it stops treating health AI as a single answer. The framework generates a synthetic patient health record, a realistic clinical vignette, and a patient agent that talks with the health AI system being.","status":"published","published_at":"2026-08-01","updated_at":"2026-08-01","record_date":"2026-08-01","date_kind":"published_at","topics":["evals","health-ai","agents","safety"],"source_urls":["https://www.amazon.science/blog/a-new-benchmark-for-evaluating-patient-facing-health-ai-agents","https://arxiv.org/abs/2607.25485"],"visuals":[{"id":"patientagentbench-health-agent-safety","kind":"editorial-diagram","role":"hero","src":"https://newruntime.com/images/posts/patientagentbench-health-agent-safety.webp","alt":"Hand-drawn benchmark pipeline where a synthetic patient record, clinical vignette, simulated patient, health agent, stateful tools, and clinician-vetted judging panel produce safety and workflow scores.","caption":"PatientAgentBench evaluates health agents where risk actually appears: sustained patient conversations, hidden clinical context, tool actions, and escalation decisions.","credit":"New Runtime synthesis from Amazon Science PatientAgentBench","source_url":"https://www.amazon.science/blog/a-new-benchmark-for-evaluating-patient-facing-health-ai-agents","generated_with":"gemini-3.1-flash-image","width":1600,"height":900,"legend":[{"label":"Synthetic patient","description":"A generated health record and clinical vignette drive a realistic multiturn conversation."},{"label":"Stateful tools","description":"The system under evaluation must reason over records, converse, and execute healthcare workflows."},{"label":"Safety jury","description":"An LLM-as-a-jury panel scores clinical safety, triage, workflow accuracy, completion, helpfulness, and conversation quality."}]}],"routes":{"html":"https://newruntime.com/posts/patientagentbench-health-agent-safety/","markdown":"https://newruntime.com/posts/patientagentbench-health-agent-safety.md","json":"https://newruntime.com/posts/patientagentbench-health-agent-safety.json"},"source_format":"markdown","next_reads":[{"type":"topic","path":"/topics/agents/","reason":"Explore the agents topic hub.","url":"https://newruntime.com/topics/agents/","title":"Agents - New Runtime","media_type":"text/html"},{"type":"topic","path":"/topics/evals/","reason":"Explore the evals topic hub.","url":"https://newruntime.com/topics/evals/","title":"Agent evals - New Runtime","media_type":"text/html"},{"type":"related_material","path":"/posts/contextual-agent-memory-four-layer-system/","reason":"Shares agents and evals.","url":"https://newruntime.com/posts/contextual-agent-memory-four-layer-system/","title":"A Vector Store Is Not An Agent Memory System","media_type":"text/html"},{"type":"related_material","path":"/posts/ramp-agentic-risk-operations/","reason":"Shares agents and evals.","url":"https://newruntime.com/posts/ramp-agentic-risk-operations/","title":"Ramp Separates Agent Reasoning From Risk Decisions","media_type":"text/html"},{"type":"related_material","path":"/posts/langchain-deep-agents-lean-harness/","reason":"Shares agents and evals.","url":"https://newruntime.com/posts/langchain-deep-agents-lean-harness/","title":"LangChain Deep Agents Shrink the Harness Instead of Adding More Prompt","media_type":"text/html"}]}
