Field note
Astra turns eval thresholds into release controls. The runtime question is when a model version can no longer ship under the previous access policy.
Why it matters
OpenAI’s own post on responding to next-frontier critical cyber capabilities is a primary-source safety/governance item with material implications for model risk thresholds, cyber capability evaluation, and deployment policy. It has concrete reader value for AI safety, labs, and security builders.
New Runtime view
Astra turns eval thresholds into release controls. The runtime question is when a model version can no longer ship under the previous access policy.
Mechanism: Cyber capability eval levels trigger mitigations and altered deployment posture.
Architectural boundary: Model capability discovery is separated from release/access authority.
Measured consequence: The threshold classification and deployment response are the measurable operational consequence; exploit details are constrained.
What remains open
- Cyber evals can miss real-world compositions.
- Public detail is limited by dual-use risk.
- Mitigation policy depends on institutional judgment.
Sources
- <https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/>