Skill
DSPy GEPA reflective optimizer
GEPA (ICLR 2026) reflective two-model prompt optimizer: local inference LM + subscription reflection LM reads failure text to propose targeted instruction rewrites
Primitives inside (8)
cheap-inference-strong-reflection-splitcalibrationSplit an optimization loop across two models by call profile: a cheap local model runs ~99% of forward passes, a strong subscription model (no per-call cost) runs the ~1% of reflection steps — the whole search runs at zero pay-per-use spend.
When: Budgeting an automated search that needs many forward passes but only occasional high-quality reasoning steps.
fixture-threshold-before-optimizationcalibrationRun automated instruction search (DSPy GEPA or any labeled-data optimizer) only once ~30-50 labeled (input, expected_output) fixtures with rubric scores exist plus ~10-15 for validation — below that, build the dataset first (the eval-dataset flywheel over existing outputs), never launch on a handful.
When: Deciding whether a module / prompt / pipeline node is ready for a labeled-data optimizer run (GEPA, instruction search).
gepa-availability-preflightfailsafeBefore building a GEPA pipeline, verify the optimizer actually exists in the installed dspy: python3 -c "import dspy; print(hasattr(dspy,'GEPA'))" — GEPA landed in dspy >= 2.6 and the exact version boundary is untested.
When: Environment setup for a GEPA run on a machine whose dspy version is unknown.
gepa-vs-grpo-rollout-economycalibrationPer the ICLR 2026 GEPA paper, reflective prompt evolution beats RL-based GRPO by up to ~20% at ~35x fewer rollouts — when rollouts are the scarce resource, choose reflective search over RL-style optimization.
When: Choosing an optimization strategy for prompt/instruction improvement under a compute budget.
metric-score-plus-feedback-shapequery-shapeA GEPA metric must return dspy.Prediction(score=<0..1>, feedback=<natural-language failure explanation>) — the feedback string is the substrate the reflection LM reads; a bare float starves the reflective loop.
When: Writing the metric function for a dspy.GEPA compile.
optimizer-family-selectioncalibrationPick the prompt-optimizer family by two questions: do you have 30-50 labeled pairs (MIPROv2/GEPA) or only a prose rubric (TextGrad/OPRO), and do you need a strong reflection model (GEPA) or must it run fully local (MIPROv2/TextGrad/OPRO)?
When: Choosing an automated prompt-optimization approach for a module given the available data and compute constraints.
pareto-frontier-variant-retentioncalibrationGEPA keeps a Pareto frontier of non-dominated program variants across rounds instead of a single best — expect the output to be a frontier from which the winner is drawn, and don't discard runners-up mid-run.
When: Interpreting and persisting GEPA optimization state/results.
reflective-failure-rewritetool-sequenceGEPA's core move: the reflection LM reads accumulated failure-feedback strings from the metric, diagnoses the pattern, and proposes a targeted instruction rewrite — mutation guided by explained failures, not random perturbation.
When: Automated instruction search where per-example failure explanations are available.
Get the whole skill
All 8 primitives of this skill as one package, with the order to apply them.
Buy only the primitives you need
Each primitive is 1 credit (≈ €0.10). Pick them from the list above — the button is next to each one.
Upgrade your own skill
Paste your skill; we pick the 5 primitives from the shelf that fit it best, as one bundle for 5 credits (≈ €0.50).
Upgrade my skillNeighbour skills
Skills whose primitives are closest to this one (bge-m3 similarity):