/Catalogue/Prompt/mohitagw15856/mohitagw15856-pm-claude-skills-ai-eval-plan

Origin: github

ai-eval-plan

Design an evaluation plan for an LLM or AI feature before shipping it. Use when asked how to evaluate a prompt/model/agent, set up an eval harness, define quality metrics for an AI feature, or build a regression gate. Produces an eval plan — task definition, datasets, metrics & rubrics, baselines, automated + human evals, a pass bar, and a regression gate.

by mohitagw15856 · updated 3h ago · imported from GitHub

Installs0+0/7d
Security score75/100
Retention 14d0%
GitHub stars1.4K

Skill logic

Execution graph
User message
Prompt rewrites behaviour
Response

SKILL.md

View on GitHub ↗

AI Eval Plan Skill

You can't improve an AI feature you can't measure, and "it looks good in the demo" is not measurement. This skill produces an evaluation plan that turns a fuzzy quality goal into a repeatable, gated test — so a prompt change that quietly makes outputs worse can't ship.

Required Inputs

Ask for these only if they aren't already provided:

  • The feature & task — what the model does and what "good output" means to a user.
  • Failure modes that matter — what bad looks like (hallucination, wrong format, unsafe, off-tone, too slow).
  • Available data — any real examples, logs, or labelled cases; or note there are none yet.
  • Who judges quality — automated checks, an LLM judge, human raters, or a mix.
  • The decision this gates — ship/no-ship, model selection, or prompt iteration.

Output Format

Eval Plan: [feature]

1. What we're measuring — the task, and a one-line definition of a good vs. bad response.

2. Eval dataset

  • Cases: how many, where they come from (real logs > synthetic), and how they're split (smoke set vs. full set).
  • Coverage: the slices/scenarios that must be represented (edge cases, adversarial, each major input type).
  • Golden answers / references: present or not, and how they were created.

3. Metrics & rubric

  • Per-dimension scores — define each dimension (e.g. correctness, grounding, format, safety, tone) on an explicit 1–5 rubric with anchor descriptions, not vibes.
  • Automated checks — deterministic assertions first (valid JSON, contains required fields, no PII, latency budget).
  • LLM-as-judge — the judge prompt, the rubric it applies, and how you guard against its bias (calibrate against human labels on a sample).
  • Human eval — when it's required (safety, subjective quality) and the rater instructions.

4. Baselines — what each candidate is compared against (current prompt, previous model, a plain-prompt control).

5. The bar — the explicit threshold to ship (e.g. "≥4.2 avg correctness, 0 safety failures, p95 < 3s") and what happens if it's missed.

6. Regression gate — how this runs in CI on every change, and the score-drop threshold that blocks a merge.

Quality Checks

  • Each metric has an explicit rubric with anchors — not just a name
  • Deterministic/automated checks are used wherever possible before reaching for an LLM judge
  • The LLM judge is calibrated against human labels on at least a sample
  • The eval set includes adversarial and edge cases, not just happy-path examples
  • There is a single, explicit numeric bar for the ship decision
  • The plan specifies how it runs as a regression gate, not just a one-time check

Anti-Patterns

  • Do not rely on a single overall score — a feature can pass on average while failing every safety case
  • Do not trust an LLM judge you haven't calibrated against humans — it has its own blind spots and biases
  • Do not eval only on happy-path inputs — the failures live in the edges and the adversarial cases
  • Do not let the eval set leak into the prompt/few-shot examples — that's training on the test set
  • Do not define the pass bar after seeing the scores — set the threshold before you run, or it means nothing

Based On

LLM evaluation practice — task-grounded rubrics, LLM-as-judge with human calibration, and regression-gated CI evals.

Discussion

No comments yet — start the thread.

Sign in to join the discussion.

/More from mohitagw15856/pm-claude-skills

mohitagw15856· 3h agoSandbox
agent-readiness-audit

Prompts · HTML · v0.1.0

Audit whether AI agents can actually use your product — docs, APIs, onboarding, errors, and discoverability, evaluated from a non-human user's perspective. Use when asked if a product is agent-ready, to audit a site or API for AI usability, to prepare for agentic traffic, or when agents keep failing against your product. Produces a scored readiness report with per-surface findings and a prioritised fix list. For optimising a single article for AI citation use aeo-optimizer; for designing the MCP server itself use mcp-server-spec.

#agent-skills#agents#ai-agents

0 1.4K
mohitagw15856· 3h agoSandbox
agm-in-a-box

Prompts · HTML · v0.1.0

Run a club, PTA, or association AGM that finishes on time and holds up later — the notice and agenda done right, a quorum plan, minutes that capture decisions not conversations, elections without awkwardness, and the follow-up that makes decisions real. Use when a volunteer says 'I have to run the AGM', 'what goes in the agenda', 'nobody comes to our meetings', or 'our elections are a mess'. Produces the notice, agenda, chair's script, minutes template, and quorum rescue plan.

#agent-skills#agents#ai-agents

0 1.4K
mohitagw15856· 3h agoSandbox
ai-feature-prd

Prompts · HTML · v0.1.0

Write a PRD for an AI-powered feature, covering the things normal PRDs miss. Use when asked to spec an AI/LLM feature, write a PRD for a feature that uses a model, or plan an AI capability (assistant, summarizer, generator, classifier). Produces an AI feature PRD — problem & UX of uncertainty, model approach, eval criteria, guardrails, fallback behaviour, the data flywheel, and cost/latency budget.

#agent-skills#agents#ai-agents

0 1.4K
mohitagw15856· 3h agoSandbox
accommodation-request

Prompts · HTML · v0.1.0

Request a reasonable accommodation at work or in education — frame it around the barrier and the adjustment (not your diagnosis), cite the right process, and navigate the back-and-forth constructively. Use when someone says 'I need a workplace accommodation', 'request reasonable adjustments', 'ADA/Equality Act accommodation', or 'how do I ask for accommodations for my disability/condition'. Produces the request letter, a barriers-and-adjustments map, disclosure guidance, and a plan for the interactive process. Not legal advice — routes to the formal process and to advocacy where needed.

#agent-skills#agents#ai-agents

0 1.4K