{"name":"OrcaReplay: Time Travel for AI Agents, Debugging and Evaluation","description":"OrcaReplay introduces \"time travel\" capabilities for AI agents, allowing developers to record, replay, fork, and debug any agent run with any model. It addresses the challenges of AI agent debugging by providing byte-for-byte reproducibility, offline analysis, and the ability to compare different models from specific checkpoints. This tool, built by the OrcaRouter.ai team, enhances observability and control over complex agent behaviors.","github":"https://github.com/Continuum-AI-Corp/OrcaReplay","url":"https://osrepos.com/repo/continuum-ai-corp-orcareplay","source":"osrepos.com","sourceDescription":"This repository profile is provided by osrepos.com, an open source repository discovery platform.","repositoryProfile":"https://osrepos.com/repo/continuum-ai-corp-orcareplay","generatedFor":"open source discovery and AI-assisted research","markdown":"https://osrepos.com/repo/continuum-ai-corp-orcareplay.md","json":"https://osrepos.com/repo/continuum-ai-corp-orcareplay.json","topics":["agent-debugging","ai-agent","llm-agents","observability","llm-evaluation","TypeScript","AI","Development Tools"],"keywords":["agent-debugging","ai-agent","llm-agents","observability","llm-evaluation","TypeScript","AI","Development Tools"],"stars":null,"summary":"OrcaReplay introduces \"time travel\" capabilities for AI agents, allowing developers to record, replay, fork, and debug any agent run with any model. It addresses the challenges of AI agent debugging by providing byte-for-byte reproducibility, offline analysis, and the ability to compare different models from specific checkpoints. This tool, built by the OrcaRouter.ai team, enhances observability and control over complex agent behaviors.","content":"## Introduction\n\nOrcaReplay, developed by the OrcaRouter.ai team, is a powerful tool designed to bring \"time travel\" capabilities to AI agents. It allows you to record, replay, fork, and debug any agent run with any model, providing unprecedented control and insight into agent behavior. Imagine your agent breaking something at 2 AM, and being able to replay it exactly, offline, as many times as you like at 9 AM, without incurring additional costs or network dependencies.\n\n## Why Use OrcaReplay and Key Benefits\n\nTraditional AI agent debugging often feels like archaeology, involving endless terminal scrolling, inconsistent re-runs, and adding print statements to unfamiliar code. Existing observability tools typically focus on costs and token usage, rather than answering the critical question: \"Why did it delete my migration file?\" OrcaReplay solves this by giving you the entire run back, complete and reproducible.\n\nKey benefits and features include:\n\n*   **Exact Reproducibility**: Record any agent and reproduce the run byte-for-byte, offline, with no model calls, tokens, or charges. This ensures consistent debugging environments.\n*   **Forking and Comparison**: From any step in a recorded run, you can fork it onto a different model and compare their outcomes. This is invaluable for evaluating model performance and identifying optimal configurations.\n*   **No Agent Modification Required**: OrcaReplay works by standing up a local proxy and setting two environment variables, capturing traffic at the process and socket boundary. This means it works with any agent, regardless of whether you own or can edit its code, or even if it uses its own API key.\n*   **Comprehensive Capture**: It sees beyond the model API, capturing shell exit codes, file writes, and other side effects that SDK wrappers cannot reach. It even records agents with no API endpoint to redirect, using TLS interception.\n*   **Persistent Traces**: Runs are saved as self-contained files, meaning you can analyze them long after the terminal session has closed.\n*   **Causal Graphs**: Visualize what produced what with `orca graph`. It shows recorded edges (direct cause-and-effect) and inferred edges (derived relationships), helping you understand complex agent interactions.\n*   **Interactive UI**: The `orca ui` command opens a self-contained HTML file, providing a visual timeline of the run. You can filter events, step through the run, or watch it play back at its actual pace.\n*   **Model Comparison with Cost Analysis**: `orca compare` forks a recorded run onto multiple models from the same checkpoint, grading each with a command you choose. It provides real token counts and costs, allowing for informed decisions.\n*   **Run Sharing**: `orca push` and `orca pull` commands facilitate sharing recorded runs between local machines and a gateway, enabling collaborative debugging and analysis.\n\n## Installation\n\nOrcaReplay requires Node.js version 20 or newer. It has no native dependencies, ensuring a smooth installation process.\n\nTo install globally:\n\nconsole\nnpm i -g orcareplay\norca doctor\n\n\n`orca doctor` checks your environment for Node, Git, and available agents.\n\n## Examples\n\n### Try it in three commands\n\nGet started quickly with `orcareplay` by installing the `orca` command and recording your first agent run:\n\nconsole\nnpm i -g orcareplay             # the package is orcareplay; the command it installs is orca\n\norca record claude              # your agent, unmodified, doing whatever it does\norca replay last                # the same run again, no network, no tokens, no charge\norca replay last --from 4 --model claude-haiku-4-5 --ui\n\n\nThe third command demonstrates the power of forking, allowing you to run the same scenario from a specific step with a different model, with the UI for visualization.\n\n### What a bug hunt actually looks like\n\nSuppose your agent was supposed to fix a failing authentication test, exited successfully, but the test still fails. OrcaReplay helps you uncover the truth:\n\nconsole\n$ orca show last\nrun_6473f858b59e  generic-openai@0.1.0  14 events  exit 0\n\nSEQ  KIND   WHAT                                            DETAIL\n0    RUN    run started                                     generic-openai\n1    SNAP   tree 919d32ba037537b43814c83779963b2cc3023db7   0 changed\n2    MODEL  claude-opus-5                                   1 messages\n3    MODEL  claude-opus-5                                   stop: tool_use · 100 in · 20 out\n4    TOOL   edit_file                                       {\"path\":\"auth.ts\",…}\n5    SNAP   tree c6af62b75c0c8b8938bd6087328b5148f3dcd534   1 changed\n6    FILE   auth.ts                                         modified +1 ?3\n7    TOOL   edit_file                                       ok\n8    MODEL  claude-opus-5                                   3 messages\n9    MODEL  claude-opus-5                                   stop: end_turn · 101 in · 5 out\n10   SNAP   tree c6af62b75c0c8b8938bd6087328b5148f3dcd534   0 changed\n11   SHELL  [\"sh\",\"-c\",\"node --check nonexistent-file.ts\"]  /tmp/hunt\n12   SHELL  shell result                                    exit 1 · 43ms\n13   RUN    run ended                                       exit 0\n\ninfo usage input=201 output=25 cost=$0.004890\n\n\nThis output reveals crucial details: the file `auth.ts` was indeed modified (seq 6), but the agent's internal check failed (seq 12, `exit 1`), yet the *agent itself* exited 0. This discrepancy is why the test still failed. You can then use `orca graph last` to visualize the causal chain of events, highlighting what led to the failure.\n\nTo compare how different models would handle the same bug, use `orca compare`:\n\nconsole\n$ orca compare last --from 5 --models claude-opus-5,claude-haiku-4-5 --verify \"npm test\"\nMODEL             VERDICT  TOKENS  COST       WALL  RUN\nclaude-opus-5     pass     201/25  $0.004890  0.3s  run_1457b35062ba\nclaude-haiku-4-5  pass     201/25  $0.000326  0.3s  run_b8ee08479fb6\n\n\nThis shows both models passed, but one cost 15 times less, providing clear data for optimization.\n\n## Links\n\n*   **GitHub Repository**: [OrcaReplay](https://github.com/Continuum-AI-Corp/OrcaReplay){:target=\"_blank\"}\n*   **Built by OrcaRouter**: [OrcaRouter.ai](https://www.orcarouter.ai){:target=\"_blank\"}\n*   **All Model APIs**: [OrcaRouter Models](https://www.orcarouter.ai/models){:target=\"_blank\"}\n*   **X (formerly Twitter)**: [@OrcaRouter](https://x.com/OrcaRouter){:target=\"_blank\"}\n*   **Discord**: [OrcaRouter Community](https://discord.com/invite/YEubt8enRA){:target=\"_blank\"}\n*   **Hugging Face**: [OrcaRouter](https://huggingface.co/orcarouter){:target=\"_blank\"}\n*   **Ollama**: [OrcaRouter](https://ollama.com/orcarouter){:target=\"_blank\"}","metrics":{"detailViews":0,"githubClicks":0},"dates":{"published":null,"modified":"2026-10-02T16:43:10.000Z"}}