GitHub Projects for AI Engineers

140 repositories that keep resurfacing in the AI engineering conversation, organized by the part of the stack they are changing.

RADAR / 2026.07 VERIFIED
Projects
140
Combined stars
5.7M
Combined forks
701.7K
Theme lanes
5

GitHub metrics checked 23 Jul 2026. Ranked by stars, then forks.

01 / SIGNAL

Popularity skyline

The ten most-starred projects in this edition. Height represents GitHub stars.

02 / ECOSYSTEM

One stack, two ways to read it

Browse by capability when choosing components, or by production depth when diagnosing what sits underneath a working AI feature. Every item opens an official repository, product page, or documentation route.

Tools
98
Capability lanes
13
GitHub links
65
Official links
33

This is a selection map, not a recommendation to use every layer or tool.

03 / DIRECTORY

Explore the stack

Use a theme lane, search by project or idea, and change the ranking signal.

140 projects

01

Tools & agent workbenches

Runtimes, coding agents, orchestration frameworks, and operator surfaces where AI work actually happens.

57
#01Personal agent runtime

openclaw

openclaw

Its rapid adoption shows demand for a persistent, cross-platform agent runtime that connects models, tools, channels, and personal workflows.

383.9K 80.7K2 export hits
#02Agent skills methodology

superpowers

obra

Reusable skills are becoming an engineering methodology: agents load bounded procedures instead of improvising every workflow from a prompt.

259.7K 23.2K
#03Learning agent runtime

hermes-agent

NousResearch

Hermes makes learning from completed work part of the runtime, turning successful procedures into durable capabilities.

219.2K 41.5K8 export hits
#05Engineering skills

skills

mattpocock

Production-tested engineering procedures packaged as skills make expert workflows portable across agent clients and teams.

184.3K 15.8K
#08Agent app platform

dify

langgenius

Dify packages workflows, RAG, tools, and deployment into a shared product surface instead of leaving teams with isolated prompts.

149.9K 23.6K
#09Agent framework

langchain

langchain-ai

LangChain remains a central abstraction layer for models, tools, retrieval, and production agent plumbing.

142.4K 23.7K
#11Coding agent

claude-code

anthropics

Terminal-native coding agents turn codebase work into a tool-executing, reviewable loop instead of a chat transcript.

138.8K 22.3K
#12Reference apps

awesome-llm-apps

Shubhamsaboo

A large runnable example index shows which RAG, agent, and workflow patterns developers actually clone and remix.

126.6K 18.7K4 export hits
#13Role-based agent stack

gstack

garrytan

Opinionated specialist roles package product, design, engineering, release, and QA work into a repeatable agent operating model.

123.8K 18.5K2 export hits
#15Browser agents

browser-use

browser-use

Browser automation is becoming a first-class capability layer for agents that inspect and act on live web apps.

106.2K 11.7K
#16Coding agent

gemini-cli

google-gemini

A model-backed CLI with tool access and MCP support fits the shift from model choice to routable development harnesses.

106.1K 14.3K
#17Coding agent

codex

openai

Codex makes local repository work goal-scoped, command-aware, and evidence-oriented.

100.8K 15.1K
#23Long-horizon agent

deer-flow

bytedance

Sandboxes, memory, tools, skills, and subagents are converging into one harness for tasks that run for minutes or hours.

77.7K 10.6K3 export hits
#24Recipes & harnesses

openai-cookbook

openai

Official examples turn API features into reusable patterns for evaluation, tool use, retrieval, and multimodal work.

74.8K 12.7K
#26Coding agent SDK

cline

cline

Cline's move across IDE, CLI, and SDK surfaces shows coding agents becoming embeddable infrastructure rather than one editor feature.

65K 7K
#28Agentic terminal

warp

warpdotdev

The terminal is becoming an operator cockpit where agents, commands, project context, and feedback loops meet.

63.6K 5.3K
#31Multi-agent orchestration

autogen

microsoft

AutoGen helped normalize agentic programs as explicit conversations, tools, and coordination patterns.

59.9K 9K
#32Multi-agent orchestration

crewAI

crewAIInc

Crew-style roles remain a practical entry point for decomposing agent work into specialist responsibilities.

56K 7.9K
#34Operations agent

goose

aaif-goose

Goose represents the open-source move from suggestions to installed, executing, tool-using desktop and code agents.

51.5K 5.7K
#35Applied recipes

claude-cookbooks

anthropics

Cookbooks turn model capabilities into runnable patterns for tool use, retrieval, multimodal work, and production integration.

49.4K 5.9K5 export hits
#39Agent platform

agno

agno-agi

Agno points toward full agent platforms with models, memory, tools, teams, and observability in one stack.

41.4K 5.7K
#40Browser automation CLI

agent-browser

vercel-labs

A browser CLI designed around agent calls makes web interaction scriptable, inspectable, and easier to embed in coding and operations loops.

39.1K 2.6K
#41Stateful agents

langgraph

langchain-ai

Graph-shaped state, checkpoints, and control flow are becoming the durable architecture behind long-running agents.

37.9K 6.4K
#44Agent configuration

awesome-copilot

github

Shared instructions, skills, and agent configurations show development practice moving from private prompting into reusable operational assets.

37K 4.6K
#45Personal agent harness

openhuman

tinyhumansai

Local-first personal memory plus fleet orchestration points toward assistants becoming durable operating environments rather than isolated chats.

35.3K 3.5K
#46Domain agent modules

financial-services

anthropics

Reference agents, connectors, and skills for financial work show how professional teams can be packaged as inspectable domain capability modules.

33.7K 5K
#47Plugin ecosystem

claude-plugins-official

anthropics

Curated plugins turn domain workflows and integrations into installable capability packages with explicit ownership.

32.5K 3.6K2 export hits
#48Codex agent teams

oh-my-codex

Yeachan-Heo

Hooks, agent teams, and operator HUDs extend a coding agent into a configurable multi-agent workbench.

32.2K 2.5K
#53Agent SDK

openai-agents-python

openai

The SDK packages tools, handoffs, tracing, and guardrails as code instead of one-off prompt structure.

28.1K 4.4K2 export hits
#56Skills package manager

skills

vercel-labs

A cross-client installer and discovery layer signals that agent skills are becoming a package ecosystem rather than copied folders.

27K 2.3K
#58Agent harness

deepagents

langchain-ai

DeepAgents packages planning, subagents, file state, and long-running behavior into a batteries-included harness.

26.7K 3.7K2 export hits
#59Local autonomous agent

agenticSeek

Fosowl

A fully local browsing and coding agent reflects demand for private autonomy without per-task hosted API dependence.

26.7K 3K
#60TypeScript agents

mastra

mastra-ai

Mastra shows the JavaScript stack absorbing agents as production application components.

26.5K 2.5K
#62Autonomous work manager

symphony

openai

Isolated implementation runs move the operator from supervising every command to managing goals, evidence, and completed work.

26.2K 2.6K2 export hits
#66Codex skills catalog

skills

openai

An official skills catalog moves reusable Codex procedures into a discoverable, versionable capability layer.

24.1K 1.6K
#68Agent terminal multiplexer

herdr

ogulcancelik

A terminal multiplexer with agent-aware states makes parallel coding sessions observable as work rather than anonymous panes.

19.9K 1.3K
#72Typed agent framework

pydantic-ai

pydantic

Typed dependencies, structured outputs, tools, and eval integrations bring agent construction closer to ordinary production Python engineering.

18.8K 2.4K
#74Multi-agent research

camel

camel-ai

CAMEL keeps the research thread alive around scaling laws, roles, societies, and cooperation among agents.

17.5K 2K2 export hits
#77Memory-backed coworker

rowboat

rowboatlabs

An open coworker that combines tools with persistent memory reflects the shift from chat assistants to accumulating work relationships.

16.7K 1.7K
#81Provider skill library

skills

google

Provider-maintained procedures for products and technologies make operational knowledge directly loadable by agents.

15.2K 1.2K
#82Agent desktop

eigent

eigent-ai

Open agent desktops are becoming operator workspaces where local tools, files, and multi-step execution meet.

14.6K 1.7K
#87Backend for coding agents

InsForge

InsForge

A backend surface designed for agentic coding lets agents provision data, auth, storage, compute, and deployment through one explicit contract.

12.4K 1.1K
#92Coding agent CLI

kimi-cli

MoonshotAI

Another capable terminal agent reinforces the trend toward model-swappable development harnesses with tools and repository access.

10.7K 1.2K
#95Embedded agent SDK

copilot-sdk

github

An SDK turns the coding agent from a destination product into a capability that other applications can embed.

10K 1.4K
#103Production templates

agent-starter-pack

GoogleCloudPlatform

Agent templates with CI/CD, evaluation, and observability compress the path from demo to governed cloud deployment.

6.5K 1.5K
#105Operational scripts

agent-scripts

steipete

Small reusable scripts make recurring maintenance and review workflows callable by agents instead of rediscovered ad hoc.

6.5K 5355 export hits
#106Parallel-agent worktrees

worktrunk

max-sixty

Agent-oriented worktree management makes isolated parallel coding sessions reproducible and easier to supervise.

6K 208
#110Agent delivery CLI

agents-cli

google

Creating, evaluating, and deploying agents through one CLI makes the lifecycle a reproducible engineering workflow.

5.3K 563
#113Coding agent CLI

mistral-vibe

mistralai

A minimal coding CLI from a model provider shows the harness layer becoming as strategic as the model endpoint.

4.7K 608
#115Embedded agent runtime

agentos

rivet-dev

A library-level runtime using isolates lets applications host bounded agents inside existing backends without a separate SaaS control plane.

4K 2002 export hits
#116Coding-agent recipes

cookbook

cursor

Official workflow recipes expose the operating patterns around coding agents, not only the editor feature set.

4K 470
#117Model skills

gemini-skills

google-gemini

Provider-maintained skills make model and SDK expertise loadable as procedures instead of copied documentation.

3.8K 3882 export hits
#118Visual engineering skills

skills

BuilderIO

Portable planning and visual-recap skills make graphical artifacts part of the coding-agent coordination loop.

3.8K 190
#120JavaScript agent SDK

openai-agents-js

openai

A JavaScript SDK brings tools, handoffs, tracing, and voice agents into the dominant web application runtime.

3.4K 868
#123Agent control plane

agentfield

Agent-Field

Treating agents as routable APIs and microservices puts memory, asynchronous execution, scaling, and audit into one control plane.

2.4K 3852 export hits
#124Coding-agent plugins

plugins

cursor

An official plugin specification makes coding-agent extensions installable and governed instead of hidden in personal configuration.

2.4K 187
#128Agent wallet toolkit

agentkit

coinbase

Wallet primitives let agents transact under explicit tool contracts, pushing identity, permissions, and audit into the agent economy.

1.3K 754
02

Context engineering & memory

The supply chain that turns documents, code, web pages, tools, and knowledge graphs into usable agent context.

37
#06Document ingestion

markitdown

microsoft

File-to-Markdown conversion is becoming the dependable ingestion layer agents need before retrieval, synthesis, and audit trails.

168.4K 12.1K
#07Web extraction

firecrawl

firecrawl

Web extraction is moving from scraping glue to an agent-ready primitive for monitoring, RAG, and newsroom intake.

154.7K 8.8K
#19Tool ecosystem

awesome-mcp-servers

punkpeye

MCP server catalogs show how quickly agent capabilities are becoming installable infrastructure.

91.2K 13.4K
#20Tool protocol

servers

modelcontextprotocol

The reference server collection makes tools and context sources portable across agent clients.

88.8K 11.3K
#27Document ingestion

docling

docling-project

Production context pipelines need structure-preserving conversion for PDFs and office documents before retrieval can be trustworthy.

63.7K 4.5K
#29Local code graph

codegraph

colbymchenry

A continuously synchronized local code graph reduces repeated repository scans and gives coding agents durable structural context.

61.9K 3.9K
#30Web and social context

Agent-Reach

Panniantong

One CLI for difficult web and social sources makes external evidence retrieval a reusable agent capability rather than per-source scraping glue.

60.1K 4.8K2 export hits
#33Recent-source research

last30days-skill

mvanhorn

A bounded skill for recent cross-platform research packages source discovery, recency, and synthesis into a repeatable evidence workflow.

53.2K 4.6K
#38Client-side code intelligence

GitNexus

abhigyanpatwari

A browser-local repository graph combines privacy, structural navigation, and Graph RAG without requiring a hosted indexing service.

44.6K 4.9K
#42Grounded extraction

langextract

google

Structured extraction with precise source grounding connects probabilistic models to auditable data pipelines.

37.8K 2.6K
#50Agent memory

cognee

topoteretes

Cognee turns memory into a self-hosted knowledge graph layer, matching New Runtime's shared-context thesis.

29.2K 2.8K
#51Temporal memory

graphiti

getzep

Temporal knowledge graphs make memory updateable, queryable, and less tied to one chat window.

29.1K 2.9K
#52Memory and context engine

supermemory

supermemoryai

A local-capable memory API separates persistent context from any one model, client, or conversation.

28.6K 2.5K
#54PDF ingestion

opendataloader-pdf

opendataloader-project

AI-ready PDF parsing is becoming a dedicated infrastructure layer because document layout and provenance cannot be recovered after naive text extraction.

27.7K 2.7K
#55Code context

repomix

yamadashy

Packing repositories into AI-readable context remains a practical bridge between large codebases and coding agents.

27.3K 1.4K
#57Tool protocol

fastmcp

PrefectHQ

FastMCP lowers the cost of turning internal APIs and scripts into agent-callable tools.

26.8K 2.2K
#63RAG pipelines

haystack

deepset-ai

Haystack keeps context engineering explicit through pipelines, retrieval, routing, memory, and observability.

26K 2.9K
#64Coding-agent memory

agentmemory

rohitg00

Shared persistent memory across coding agents reduces repeated onboarding and makes context continuity measurable against real tasks.

25.7K 2.1K
#69Context-window optimization

context-mode

mksglu

Sandboxing tool output and routing context through MCP and hooks turns token pressure into an explicit systems optimization problem.

19.2K 1.4K
#70Document linearization

olmocr

allenai

High-quality PDF linearization improves the training and retrieval substrate before an LLM ever sees the document.

19.2K 1.6K
#73Long-document OCR

Unlimited-OCR

baidu

One-shot long-horizon document parsing challenges page-by-page pipelines and reduces the reconstruction burden in document context systems.

18.1K 1.7K
#79Production memory

Memori

MemoriLabs

Agent-native memory that writes structured state into existing data infrastructure matches the move from chat history to governed organizational memory.

15.6K 2.9K
#83Cross-agent memory

memU

NevaMind-AI

Memory shared across agents addresses the fragmentation caused when every client maintains a separate understanding of the user.

14.1K 1K
#85Codebase documentation

openwiki

langchain-ai

Agent-maintained repository documentation creates a durable context layer that stays closer to the changing code.

13K 8962 export hits
#86Local private RAG

LEANN

StarTrail-org

Storage-efficient local retrieval makes private personal and edge-device RAG practical without a heavyweight vector database.

12.7K 1.1K
#88Code search MCP

claude-context

zilliztech

Repository-wide semantic search exposed through MCP gives multiple coding clients a shared context service.

12.2K 904
#90Incremental context

cocoindex

cocoindex-io

Incremental indexing keeps long-running agent context fresh without paying for full reprocessing after every source change.

11K 848
#98Repository context

git-mcp

idosal

A remote MCP context layer gives agents source-grounded access to repositories without manually packing every codebase into the prompt.

8.3K 735
#102Visual context compression

pxpipe

teamchong

Rendering bulky text context as images explores a new tradeoff between multimodal input capacity, token cost, and recoverability.

6.6K 5652 export hits
#104Context retrieval

airweave

airweave-ai

A dedicated retrieval layer separates continuously changing source synchronization from the agent application itself.

6.5K 816
#108Code search

semble

MinishLab

Token-efficient structural code search reduces the context tax of finding the right repository evidence before editing.

5.7K 237
#114Multimodal search

mgrep

mixedbread-ai

Semantic grep across code, images, and documents offers agents one retrieval primitive across heterogeneous project context.

4.3K 171
#119Skill memory

Acontext

memodb-io

Treating skills as memory connects reusable procedures with the state and evidence agents need to apply them correctly.

3.6K 326
#121Agent storage

seekdb

oceanbase

Unifying vector, text, structured, and semi-structured retrieval reduces the storage fragmentation behind agent memory.

2.8K 311
#122File context packing

files-to-prompt

simonw

A small deterministic context packer remains useful because inspectable preprocessing is often preferable to opaque ingestion.

2.8K 176
#127Local code intelligence

chunkhound

chunkhound

Local-first codebase intelligence gives agents structural retrieval without exporting private repositories to a hosted index.

1.4K 114
#132Enterprise work context

work-iq

microsoft

An MCP and CLI surface for organizational work context points toward governed enterprise knowledge becoming directly agent-callable.

947 109
03

LLM-UI & generated interfaces

Interfaces assembled around model output, task state, validation, and human correction.

14
#36Programmatic text layout

pretext

chenglou

Fast deterministic text measurement gives generative interfaces a layout primitive agents can calculate and render instead of approximating.

49.3K 2.7K
#37Design language for agents

impeccable

pbakaus

An explicit design language helps coding agents reason about visual tradeoffs instead of reproducing generic component patterns.

49K 2.9K
#43Agent-native video

hyperframes

heygen-com

HTML-to-video gives agents a code-native visual medium they can generate, inspect, revise, and render deterministically.

37K 3.5K
#61Design context standard

design.md

google-labs-code

A portable design-system file gives coding agents durable visual context instead of forcing them to infer taste from screenshots on every task.

26.3K 2.1K
#78Generative UI

json-render

vercel-labs

Structured UI output lets model answers become inspectable interfaces instead of unstructured prose blobs.

15.7K 849
#89Generative UI

tambo

tambo-ai

A React generative-UI SDK points to applications where the interface is assembled around the task state.

11.2K 559
#96Agent interface

magentic-ui

microsoft

Browser and local-file agents need an operator interface, not just a backend chain of tool calls.

10K 1K
#100Design agent skills

stitch-skills

google-labs-code

Design procedures packaged as portable skills connect generative UI tools to repeatable production workflows.

7.8K 916
#107Agent-native slides

open-slide

1weiho

A code-native slide framework gives agents structured layout, components, and renderable output rather than opaque presentation files.

6K 416
#112Streaming diagram UI

drawio-mcp

jgraph

Streaming diagram primitives through MCP turns model-generated structure into an inspectable visual work surface while it is being built.

4.9K 3102 export hits
#125Embedded agent UI

openai-apps-sdk-examples

openai

Runnable Apps SDK examples show how model responses become interactive, stateful product surfaces instead of plain text.

2.3K 516
#133Realtime voice UI

realtime-voice-component

openai

A reusable voice component makes low-latency multimodal interaction an embeddable interface primitive rather than a bespoke demo.

878 113
#134Generative UI protocols

generative-ui

CopilotKit

Examples across AG-UI, A2UI, Open JSON UI, and MCP Apps reveal an emerging protocol layer for model-generated interfaces.

783 69
#136Design reasoning skill

taste-skill

senlindesign

Extracting design decisions and tradeoffs rather than only tokens gives interface agents a more durable representation of visual intent.

241 17
04

Evals, specs & reliability

The control plane: specifications, tests, protocols, structured output, security boundaries, and sandboxes.

28
#04Coding behavior contract

andrej-karpathy-skills

multica-ai

A compact CLAUDE.md derived from recurring coding-agent failures shows how operational lessons are becoming reusable behavioral contracts.

195.7K 20.1K2 export hits
#10Harness transparency

system-prompts-and-models-of-ai-tools

x1xhlol

The collection exposes how production AI tools combine prompts, tool contracts, and hidden orchestration, making harness design inspectable.

142.2K 34.8K5 export hits
#14Specification workflow

spec-kit

github

Spec-driven development makes requirements, plans, and acceptance checks the control surface before agents write code.

123.4K 11K
#18Autonomous research loop

autoresearch

karpathy

A bounded single-GPU research loop turns hypothesis, experiment, measurement, and iteration into an inspectable autonomous workflow.

91.9K 13.1K
#21Agent code restraint

ponytail

DietrichGebert

Encoding senior-engineer restraint into agent behavior addresses the growing cost of unnecessary abstractions and generated-code sprawl.

88.3K 4.8K
#65Agent protocol

A2A

a2aproject

Agent-to-agent protocols matter as soon as agents become opaque services owned by different products or teams.

25K 2.5K
#67Agent skills standard

agentskills

agentskills

A shared skill specification makes procedures portable across agent clients and gives teams a stable contract for capability packaging.

23.4K 1.6K
#71Evaluation

evals

openai

Evals are moving from model benchmarks into product infrastructure for judging complete AI systems.

19K 3K
#75Agent optimization

agent-lightning

microsoft

Agent training is shifting from prompt tweaking toward trace-level optimization of complete trajectories and tool decisions.

17.4K 1.5K
#76Evaluation

deepeval

confident-ai

DeepEval makes LLM and agent quality measurable in CI-style loops instead of relying on manual prompt taste.

17.1K 1.7K
#80Structured generation

outlines

dottxt-ai

Grammar and schema-constrained outputs reduce the gap between probabilistic generation and software contracts.

15.2K 807
#84Skill security

SkillSpector

NVIDIA

As skills become executable supply-chain artifacts, static inspection for malicious instructions and risky behavior becomes mandatory.

13.6K 1.1K
#93Agent sandbox

CubeSandbox

TencentCloud

Isolated execution environments are required before agents can safely run code, browsers, or external tools.

10.6K 946
#94Agent security

hexstrike-ai

0x4m4

Offensive MCP tooling signals that agent tool access is both a capability frontier and a security boundary.

10.4K 2.2K
#97Model-tiered code audit

improve

shadcn

Separating expensive architectural review from cheaper execution models turns model routing into a concrete engineering workflow.

8.6K 3722 export hits
#99Secure agent interpreter

monty

pydantic

A minimal Rust-based Python interpreter gives agents a deliberately bounded execution surface instead of a full host runtime.

7.9K 394
#101LLM observability

evidently

evidentlyai

Evaluation, testing, and monitoring are converging into one operational loop for AI systems after deployment.

7.7K 886
#109Integrated agent sandbox

sandbox

agent-infra

Combining browser, shell, files, MCP, and an editor in one container gives agent runs a bounded and reproducible execution environment.

5.5K 491
#111Agent programming language

zerolang

vercel-labs

A language built around structured diagnostics gives coding agents machine-readable repair signals instead of prose compiler output.

5.2K 341
#126Skill evaluation

skillsbench

benchflow-ai

Skills need behavioral benchmarks that measure both the procedure and the agent's ability to invoke it correctly.

1.6K 343
#129Evaluation library

openevals

langchain-ai

Ready-made evaluators lower the cost of putting repeatable quality checks around LLM and agent behavior.

1.1K 106
#130Loop engineering skills

skills

AI-Builder-Club

Codebase maps and loop-engineering procedures package structural evidence and repeatable maintenance cycles for coding agents.

1K 1402 export hits
#131Agent runtime security

ClawKeeper

SafeAI-Lab-X

Skills, plugins, and runtime watchers form a layered security boundary around a tool-using personal agent.

1K 58
#135Agent service contract

agent-protocol

langchain-ai

A framework-neutral service contract separates agent infrastructure and lifecycle operations from any one orchestration library.

637 55
#137Business-workflow benchmark

AutomationBench

zapier

Realistic business automation tasks test whether agents can complete multi-step work across tools, not merely answer benchmark questions.

141 16
#138Agent memory benchmark

MemoryData

OpenDataBox

A unified benchmark for memory-augmented agents makes storage, extraction, retrieval, and use comparable as a system.

120 19
#139Pipeline self-optimization

fully-automated-prompt-optimization

cisco-foundation-ai

Optimizing multi-step LLM chains with coding agents and evals moves improvement from manual prompt editing to a measured loop.

99 12
#140Iterative agent improvement

altk-evolve

AgentToolkit

Iteration-based self-improvement makes agent changes explicit artifacts that can be compared and evaluated between runs.

95 11
05

Model & serving infrastructure

Serving throughput, local training, cache economics, and the compute layer underneath agent systems.

4
#22Model serving

vllm

vllm-project

High-throughput serving keeps model routing and self-hosted inference economically plausible.

86.9K 19.7K
#25Local models

unsloth

unslothai

Faster local training and inference give teams a cheaper private lane for model experiments.

68.8K 6.2K
#49Agent browser engine

browser

lightpanda-io

A headless browser built for automation treats browser execution cost and throughput as first-class agent infrastructure.

32.1K 1.4K
#91KV-cache

LMCache

LMCache

KV-cache infrastructure turns long context and repeated workloads into a serving optimization problem.

10.8K 1.6K
Proof / reproducibility

What has to reproduce

The JSON route exposes the ranked repository dataset, proof contract, handoff routes, and public GitHub source URLs.

Verification
locally reproduced
Outputs
6 fixtures
Replay
2 receipts
Axes
7 assessed

Fixture outputs

  • Structured radar datasetdist/projects/github-projects-for-ai-engineers.json

    The JSON route exposes the ranked repository dataset, proof contract, handoff routes, and public GitHub source URLs.

    Open route
  • Human radar pagedist/projects/github-projects-for-ai-engineers/index.html

    The HTML radar renders the same project count, theme lanes, and machine-readable handoff links as the dataset.

    Open route
  • Agent build briefdist/projects/github-projects-for-ai-engineers/build-brief.md

    The build brief names the data source, GitHub metric refresh boundary, privacy constraints, and acceptance checks.

    Open route
  • Clean-room replay fixturescripts/project-proof-depth-v2-fixture-test.mjs

    The fixture replays the bounded deterministic-ranking task in a temporary directory, proves the initial failure of the wrong sort order, proves the final pass of the stars-then-forks rule, and verifies that an unrelated file is unchanged.

  • Independent replay kitscripts/project-proof-independent-replay-kit.mjs

    The kit starts with a ranking function that violates the published sort rule, exposes no ready solution, protects every out-of-scope file by checksum, and refuses to emit a receipt without a non-author attestation and a passing final check.

  • Independent agent replay receiptsrc/data/project-proof-receipts/github-projects-for-ai-engineers-independent-agent-v1.json

    A separately spawned agent received only the isolated kit, observed the failing deterministic-ranking check, changed only src/rank.mjs, passed the final check, and produced an attestation that the primary verifier accepted.

    Open route

Replay commands

  1. npm run project-proof:replay:test
  2. npm run project-proof:replay:prepare -- --project=github-projects-for-ai-engineers
  3. npm run validate:content
  4. npm run build
  5. npm run validate:agent-indexes

Provider scopes

  • GitHubread

    Public repository metadata and repository pages only.

    Stars, forks, archived state, redirects, and canonical full names drift over time.

    Human approves a live GitHub refresh before network calls.

  • Local Telegram exportread

    Repository-link roots and duplicate counts from an approved local export snapshot.

    The radar starts from private discovery provenance but publishes only public GitHub repository URLs.

    Human names the export snapshot or existing dataset before extraction.

  • Public sitepublish

    Project content, project dataset, and generated static routes for this radar only.

    The public radar must expose the refreshed structured dataset and human visualization together.

    Owner approves commit, push, and deploy separately from the metric refresh.

Seven-axis Buildability

non author replayed · 2026-08-15

timehigh
The clean-room ranking task reaches a checked result in one bounded local run, while the public blueprint targets one local content pass plus a GitHub metric refresh.
BottleneckA full radar refresh still depends on fetching current GitHub metrics for all 140 repositories.
code burdenmedium
The replay needs one focused comparator edit plus the existing test, and the public blueprint needs normalization scripts and content validation rather than a bespoke application.
BottleneckCurating theme lanes and relevance still needs an editor who knows the AI-engineering landscape.
integration burdenmedium
The replay uses only repository files, Node, and a test runner, while the wider blueprint touches public GitHub reads and Astro content routes that already exist in this repository.
BottleneckGitHub metric drift forces every refresh through an approved external read before the ledger is current.
operational burdenhigh
The first proof is local and temporary, requires no server, and touches no shared or production state.
BottleneckKeeping the published radar current requires a recurring refresh protocol rather than one-off runs.
permission clarityhigh
GitHub read, local Telegram-export read, and public-site publish scopes are named separately with explicit approval gates in the build brief and acceptance contract.
BottleneckA live metric refresh and a public deploy still need two separate owner approvals.
reproducibilityhigh
A separately spawned agent received only the isolated public kit, reproduced the failing deterministic-ranking check, changed the one allowed comparator file, and passed the same acceptance check without inspecting the author solution.
BottleneckThe bounded fixture proves the published sort rule, not a full 140-repository refresh against live GitHub data.
failure recoveryhigh
The replay records the initial failing check, preserves an unrelated file byte-for-byte, performs no remote action, and leaves explicit residual-risk and approval sections.
BottleneckRolling back a bad public radar release is outside this local proof.

Replay evidence

Bounded receipts, with limitations kept visible.

clean-room-bounded-edit-v1passed

automated clean room · 2026-08-15

npm run project-proof:replay:test
  • This replay proves the deterministic ranking contract on a bounded fixture dataset, not a real external repository integration or a live GitHub metric refresh.
  • No human or separately operated agent independently interpreted the brief in this clean-room replay.
independent-claude-replay-v1passed

independent agent · 2026-08-15

node --test test/rank.test.mjs
  • This proves a bounded non-author replay of the public workflow contract, not integration into a production repository.
  • Reviewer identity is a local agent attestation and is not cryptographically verified.
Known limits
  • GitHub stars and forks are time-sensitive; the public method must keep the refresh date visible.
  • The first pass does not clone, run, or security-audit the listed repositories.
  • Export-hit counts are discovery signals, not public provenance; raw Telegram data remains private.
  • The ecosystem map is a navigation and architecture aid, not a compatibility matrix or recommendation to adopt every tool.
  • The bounded replay proves the deterministic ranking rule on a small fixture dataset, not a full 140-repository refresh against live GitHub metrics.
  • The independent replay proves the bounded public workflow contract, not integration into a real external or production repository.
  • The independent reviewer identity is a local agent attestation and is not cryptographically verified.

Retrieval answer

A Telegram-corpus-derived GitHub radar plus a source-backed map of the models, orchestration, retrieval, memory, security, automation, and infrastructure layers around production AI systems. Turn the Telegram export and two public ecosystem references into a ranked repository radar and a navigable map of the wider production AI stack.

Discovery graph / next reads

Continue through New Runtime

Open the graph
  1. 01related materialClient-local memory -> shared context infrastructureShift this project is designed to test.
  2. 02related materialUnder the Hood: AI Engineers Need Mechanism MapsField Note supplying context for this project.
  3. 03related materialCognee Packages Agent Memory as a Self-Hosted Knowledge GraphField Note supplying context for this project.
  4. 04related materialLong-Running Loops Need Goals, Not Keep-Going PromptsField Note supplying context for this project.
  5. 05related materialAgent APIs Need Fewer Magic Tricks and More FactsField Note supplying context for this project.

These links are also published in this page’s JSON twin and as typed edges in DiscoveryGraph v1.

Who read this page?Machine requests, hidden until opened

Loading the privacy-safe route aggregate…

Open the JSON contract