{"name":"skill-up: Evaluate and Improve Agent Skills","description":"skill-up is a Go CLI for testing Agent Skills, agents, and workspaces with repeatable cases and structured reports. It suits teams that want to catch behavior regressions or iteratively improve Skills using evaluation results.","github":"https://github.com/alibaba/skill-up","url":"https://osrepos.com/repo/alibaba-skill-up","source":"osrepos.com","sourceDescription":"This repository profile is provided by osrepos.com, an open source repository discovery platform.","repositoryProfile":"https://osrepos.com/repo/alibaba-skill-up","generatedFor":"open source discovery and AI-assisted research","markdown":"https://osrepos.com/repo/alibaba-skill-up.md","json":"https://osrepos.com/repo/alibaba-skill-up.json","topics":["go","cli","ai-agents","agent-skills","evaluation","agent-evaluation"],"keywords":["go","cli","ai-agents","agent-skills","evaluation","agent-evaluation"],"stars":null,"summary":"skill-up is a Go CLI for testing Agent Skills, agents, and workspaces with repeatable cases and structured reports. It suits teams that want to catch behavior regressions or iteratively improve Skills using evaluation results.","content":"## Overview\n\nskill-up runs declarative test cases against supported agent engines, evaluates responses and workspace changes, and produces reports for local development or CI. It addresses the difficulty of checking whether an Agent Skill or agent behaves reliably across repeatable prompts and project setups.\n\nIt also supports an improvement loop: use its bundled skill-upper workflow to review failures, adjust a Skill or its evaluation cases, and run the suite again. It is aimed at developers building Agent Skills, agent workflows, or workspace-based coding tasks who want evidence to guide iteration.\n\n## Key Features\n\n- Define evaluation cases and configuration in YAML, then validate and run them from the CLI.\n- Compare Skill-enabled runs with runs that omit the Skill, or evaluate an agent on its own.\n- Test workspace tasks using repository fixtures, supplied files, or an existing local workspace.\n- Run cases with built-in engines including Claude Code, Codex, and Qoder CLI, or configure a custom engine.\n- Grade results with rule-based checks, scripts, or an agent judge.\n- Generate JSON, JUnit XML, HTML, and benchmark reports for review or CI.\n- Import Anthropic-compatible `evals.json` files into the YAML case format.\n- Run evaluations in GitHub Actions using the repository's action.\n\n## Use Cases\n\n- **Agent Skill authors** can check whether a Skill improves results by running cases both with and without it.\n- **Coding-agent teams** can test repository tasks against fixtures and detect regressions in expected responses or file changes.\n- **Developer-tool maintainers** can compare supported agent engines against the same case suite.\n- **CI owners** can run Skill evaluations on pull requests and retain machine-readable reports.\n- **Teams iterating on Skills** can use skill-upper to interpret failed cases, update coverage, and repeat the evaluation loop.\n\n## Our Take\n\nskill-up is a young but actively maintained evaluation CLI, with frequent recent commits and releases, contributors beyond its most active author, and quick issue and pull-request handling.\n\n**Good fit if:**\n- You need repeatable evaluations for Agent Skills, agents, or coding tasks that change a workspace.\n- Your workflow uses a supported engine, or you can provide a custom engine using its documented local transport.\n- You want a Go-based CLI and are comfortable adopting a project whose interfaces and workflows may still evolve.\n\n**Look elsewhere if:**\n- You need an evaluation product that does not depend on agent-engine setup or YAML case configuration.\n- Your environment cannot use the supported engines or provide a compatible custom engine.\n- You need a mature, long-established tool rather than a project that is less than six months old and changing quickly.\n\n## Project Health\n\n| Signal | Value |\n|---|---|\n| Status | **New**: less than six months old and changing fast |\n| Last commit | 2026-09-29 (8 days ago) |\n| Commits, last 90 days | 100+ |\n| Releases, last 12 months | 19 (latest v0.12.0, 2026-09-18) |\n| Contributors | 14 (top contributor: 40% of commits) |\n| Issues closed, last 90 days | 17+ (typically closed in 2 days) |\n| Pull requests merged, last 90 days | 91 (typically merged in 1 day) |\n| Project age | 5 months |\n\nChecked on 2026-10-07 with the GitHub API.\n\n## Project Facts\n\n- Language: Go\n- License: Apache-2.0\n- Stars: 1.1k\n- Forks: 95\n- Topics: agent-skills, ai, ai-agents, alibaba, skills\n- Archived: no\n\n## What You Need\n\nDetected in the repository:\n\n- Node.js (from package.json)\n- Go 1.25.0 or newer (from go.mod)\n- A test suite and automated checks on GitHub Actions\n\n## Getting Started\n\nInstall the CLI with the repository's install script:\n\n```bash\ncurl -fsSL https://raw.githubusercontent.com/alibaba/skill-up/main/install.sh | bash\n```\n\nThen validate and run an evaluation configuration with `skill-up validate` and `skill-up run`. See the [README](https://github.com/alibaba/skill-up) and [user manual](https://alibaba.github.io/skill-up/) for setup and case-writing details.\n\n## License in Plain Words\n\n**Apache-2.0** (permissive).\n\n- Commercial use: yes\n- Modify and redistribute: yes\n- You must keep: the license, the NOTICE file and a note of your changes\n- Share your changes: no\n- Includes an explicit patent grant from the contributors\n\nA summary, not legal advice: the LICENSE file is what applies.\n\n## Alternatives\n\n- [agent-anvil](https://osrepos.com/repo/agent-axiom-agent-anvil): Agent Anvil evaluates tool-using agents with scenario runs and trace checks, while skill-up also targets Skills and workspaces with repeatable cases.\n- [agentevals](https://osrepos.com/repo/agentevals-dev-agentevals): agentevals scores existing OpenTelemetry traces without rerunning agents, while skill-up runs repeatable tests and produces structured reports.\n- [SkillOpt](https://osrepos.com/repo/microsoft-skillopt): SkillOpt uses scored task trajectories to optimize and validate skill documents, while skill-up focuses on repeatable tests and regression reports.\n\n| Project | Language | License | Stars | Status |\n|---|---|---|---|---|\n| **skill-up** | Go | Apache-2.0 | 1.1k | New |\n| [agent-anvil](https://osrepos.com/repo/agent-axiom-agent-anvil) | Python | MIT | 9 | Not checked yet |\n| [agentevals](https://osrepos.com/repo/agentevals-dev-agentevals) | Python | Apache-2.0 | 162 | Not checked yet |\n| [SkillOpt](https://osrepos.com/repo/microsoft-skillopt) | Python | MIT | 18k | New |\n\n## Considerations\n\nThe project is new and changing quickly, so teams should review releases and verify that the current engine integrations and evaluation formats fit their workflow. Its recent activity is strong, and development is not concentrated in a single contributor, but its short history means less time for interfaces and practices to settle.\n\nThe CLI requires Go 1.25 to build from source; the repository also lists Node.js as a runtime requirement. Agent evaluation additionally depends on configuring an Agent Engine and, where applicable, model credentials. The GitHub Action runs in a Linux Docker container. The project is Apache-2.0 licensed, so redistributed copies should retain the applicable license and notices.","metrics":{"detailViews":4,"githubClicks":3},"dates":{"published":null,"modified":"2026-10-07T12:25:30.000Z"}}