---
type: "post"
stable_id: "post:research-and-eval-primitives-memory-removal-tutoring-and-refactoring"
slug: "research-and-eval-primitives-memory-removal-tutoring-and-refactoring"
title: "Research and Eval Primitives: Memory, Removal, Tutoring, and Refactoring"
description: "Four narrow releases split evaluation into testable units: transferable agent memory, perceptual object-removal quality, a concrete tutoring moment, and large-scale multilingual refactoring."
retrieval_nugget: "The set is useful precisely because the metrics cannot be collapsed into one score. It shows evaluation moving from a headline leaderboard toward task-specific diagnostics."
published_at: "2026-08-14"
updated_at: "2026-08-15"
record_date: "2026-08-14"
date_kind: "discovered_at"
topics: ["research","evals","benchmarks"]
entities: ["Agent Memory Distillation","PROVE","TutorMoments","SWE-Bench ProMax"]
source_urls: ["https://agent-memory-distillation.github.io/","https://arxiv.org/abs/2605.14534","https://github.com/xiaomi-research/prove","https://allenai.org/blog/tutormoments","https://arxiv.org/abs/2608.09802"]
source_format: "article"
editorial_timing: {"lane":"regular_hourly","scheduled_at":"2026-08-21T15:00:00+03:00","real_news_delta":"owner-selected cross-source synthesis"}
origin: {"basket_id":"5854f7b2-5954-4d2a-8997-81596a49da74","basket_revision":1,"target_kind":"synthesis","target_id":"219dd943-74af-43d0-adaa-273be7065a40","owner_selection":"40","route":"openclaw"}
visual_decision: {"outcome":"text_only","status":"not_applicable","reason_code":"concise_text_sufficient","owner_reviewed":true,"reviewed_by":"owner-and-codex"}
schema_version: "newruntime-agent-readable-v0.2"
status: "published"
visuals: []
routes: {"html":"https://newruntime.com/posts/research-and-eval-primitives-memory-removal-tutoring-and-refactoring/","markdown":"https://newruntime.com/posts/research-and-eval-primitives-memory-removal-tutoring-and-refactoring.md","json":"https://newruntime.com/posts/research-and-eval-primitives-memory-removal-tutoring-and-refactoring.json"}
---

# Research and Eval Primitives: Memory, Removal, Tutoring, and Refactoring

## Retrieval answer

The set is useful precisely because the metrics cannot be collapsed into one score. It shows evaluation moving from a headline leaderboard toward task-specific diagnostics.

Four narrow releases split evaluation into testable units: transferable agent memory, perceptual object-removal quality, a concrete tutoring moment, and large-scale multilingual refactoring.

New Runtime reading: The set is useful precisely because the metrics cannot be collapsed into one score. It shows evaluation moving from a headline leaderboard toward task-specific diagnostics.

Evidence boundary: this item uses the listed public sources and keeps vendor, author, or reporter claims attributed. The queued page is an editorial synthesis, not an independent validation of every reported metric.
