Integration test quality
Claude-assisted review of environment integration tests.
Purpose
ct envs integration-test-quality uses Claude Code on the host to judge whether an environment's integration tests are truly end-to-end from a user’s perspective, with optional passes for task conflicts and coverage gaps. It writes integration_test_quality.md under the output directory.
It requires Claude Code (the claude CLI) on your PATH.
Commands
ct envs integration-test-quality
Analyze integration test quality for an environment.
--check-tasks adds a pass for whether completing main or side tasks would break the tests, --coverage adds a coverage-gap pass, and --fix then runs Claude to edit the tests from the saved report, asking for confirmation first in an interactive terminal.
Examples:
ct envs integration-test-quality -e web_scraping ct envs integration-test-quality -e web_scraping --check-tasks ct envs integration-test-quality -e web_scraping --coverage ct envs integration-test-quality -e web_scraping --check-tasks --coverage ct envs integration-test-quality -e web_scraping --fix
Output
The command writes or appends to data/test_quality_runs/<env>/integration_test_quality.md (or the file under --output-dir), prints a short summary, then prints the report path.
The main analysis is plain text with no markdown headings beyond the top title:
- A title line:
# Integration test quality: <env> Summary:, a blank line, then a 1–2 sentence overall assessment of the suite.List of problems, a blank line, then one line per issue asproblem: how to fix.
--check-tasks appends a Task conflicts section in the same shape: a header line, a blank line, then one line per task or area and why tests would break.
--coverage appends a Coverage gaps section in the same shape, with one line per gap and its suggested coverage.
Criteria (good vs bad integration tests)
Good tests:
- Exercise the system through HTTP, CLI, or other stable external interfaces
- Use inputs and read outputs the way a real user would
- Survive internal refactors because they only depend on public behavior
Bad tests:
- Import internal modules or query the DB directly instead of the product API
- Assert on internal paths, logs, or implementation details
- Over-mock so they no longer test real integration