alibaba

skill-up: Evaluate and Improve Agent Skills

skill-up is a Go CLI for testing Agent Skills, agents, and workspaces with repeatable cases and structured reports. It suits teams that want to catch behavior regressions or iteratively improve Skills using evaluation results.

Stars
1.1k
Forks
95
Last commit
8 days ago
Contributors
14
Releases, 12 months
19

Our take

Newless than six months old and changing fast
  • Merges more pull requests than 95% of the projects we track
  • Releases more often than 75% of the projects we track

skill-up is a young but actively maintained evaluation CLI, with frequent recent commits and releases, contributors beyond its most active author, and quick issue and pull-request handling.

Good fit if

  • You need repeatable evaluations for Agent Skills, agents, or coding tasks that change a workspace.
  • Your workflow uses a supported engine, or you can provide a custom engine using its documented local transport.
  • You want a Go-based CLI and are comfortable adopting a project whose interfaces and workflows may still evolve.

Look elsewhere if

  • You need an evaluation product that does not depend on agent-engine setup or YAML case configuration.
  • Your environment cannot use the supported engines or provide a compatible custom engine.
  • You need a mature, long-established tool rather than a project that is less than six months old and changing quickly.
All health signals
Last commit2026-09-29 (8 days ago)
Commits, last 90 days100+
Releases, last 12 months19 (latest v0.12.0, 2026-09-18)
Contributors14 (top contributor: 40% of commits)
Issues closed, last 90 days17+ (typically closed in 2 days)
Pull requests merged, last 90 days91 (typically merged in 1 day)
Project age5 months

Checked on 2026-10-07 with the GitHub API.

Overview

skill-up runs declarative test cases against supported agent engines, evaluates responses and workspace changes, and produces reports for local development or CI. It addresses the difficulty of checking whether an Agent Skill or agent behaves reliably across repeatable prompts and project setups.

It also supports an improvement loop: use its bundled skill-upper workflow to review failures, adjust a Skill or its evaluation cases, and run the suite again. It is aimed at developers building Agent Skills, agent workflows, or workspace-based coding tasks who want evidence to guide iteration.

Key Features

  • Define evaluation cases and configuration in YAML, then validate and run them from the CLI.
  • Compare Skill-enabled runs with runs that omit the Skill, or evaluate an agent on its own.
  • Test workspace tasks using repository fixtures, supplied files, or an existing local workspace.
  • Run cases with built-in engines including Claude Code, Codex, and Qoder CLI, or configure a custom engine.
  • Grade results with rule-based checks, scripts, or an agent judge.
  • Generate JSON, JUnit XML, HTML, and benchmark reports for review or CI.
  • Import Anthropic-compatible evals.json files into the YAML case format.
  • Run evaluations in GitHub Actions using the repository's action.

Use Cases

  • Agent Skill authors can check whether a Skill improves results by running cases both with and without it.
  • Coding-agent teams can test repository tasks against fixtures and detect regressions in expected responses or file changes.
  • Developer-tool maintainers can compare supported agent engines against the same case suite.
  • CI owners can run Skill evaluations on pull requests and retain machine-readable reports.
  • Teams iterating on Skills can use skill-upper to interpret failed cases, update coverage, and repeat the evaluation loop.

What you need

Detected in the repository

  • Node.js (from package.json)
  • Go 1.25.0 or newer (from go.mod)
  • A test suite and automated checks on GitHub Actions

License in plain words

Apache-2.0permissive

  • Commercial use: yes
  • Modify and redistribute: yes
  • You must keep: the license, the NOTICE file and a note of your changes
  • Share your changes: no
  • Includes an explicit patent grant from the contributors:

A summary, not legal advice: the LICENSE file is what applies.

Getting Started

Install the CLI with the repository's install script:

curl -fsSL https://raw.githubusercontent.com/alibaba/skill-up/main/install.sh | bash

Then validate and run an evaluation configuration with skill-up validate and skill-up run. See the README and user manual for setup and case-writing details.

Alternatives

  • agent-anvil: Agent Anvil evaluates tool-using agents with scenario runs and trace checks, while skill-up also targets Skills and workspaces with repeatable cases.
  • agentevals: agentevals scores existing OpenTelemetry traces without rerunning agents, while skill-up runs repeatable tests and produces structured reports.
  • SkillOpt: SkillOpt uses scored task trajectories to optimize and validate skill documents, while skill-up focuses on repeatable tests and regression reports.
ProjectLanguageLicenseStarsStatus
skill-upGoApache-2.01.1kNew
agent-anvilPythonMIT9New
agentevalsPythonApache-2.0162Active
SkillOptPythonMIT18kNew

Considerations

The project is new and changing quickly, so teams should review releases and verify that the current engine integrations and evaluation formats fit their workflow. Its recent activity is strong, and development is not concentrated in a single contributor, but its short history means less time for interfaces and practices to settle.

The CLI requires Go 1.25 to build from source; the repository also lists Node.js as a runtime requirement. Agent evaluation additionally depends on configuring an Agent Engine and, where applicable, model credentials. The GitHub Action runs in a Linux Docker container. The project is Apache-2.0 licensed, so redistributed copies should retain the applicable license and notices.

Found this useful?

Share it with someone who would like skill-up.

Comparisons

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️