Skill
Context pipeline assembly
6-stage context assembly pipeline (Collect→Filter→Compress→Arrange→Format→Assemble) that curates what enters an LLM context window before every model call; the critical missing Arrange step places critical content at START+END for attention distribution
Primitives inside (8)
assemble-budget-check-loopbackfailsafeAt final context assembly, validate total tokens against the model's limit; if over budget, loop back to stronger compression or stricter filtering — never truncate blindly or ship an over-limit prompt.
When: The assembled context string/messages array is ready to send to the model.
explicit-section-delimitersdisciplineWrap each context section in explicit delimiters (XML tags per section, consistent JSON schemas for tool outputs, markdown inside sections) and put few-shot examples in a dedicated block at the high-attention END, so the model always knows which section is which.
When: A multi-section prompt (system, history, retrieved docs, tool outputs, question) is being formatted for the model.
filter-per-step-not-per-taskcalibrationJudge context relevance against the current step at each assembly, drawing from the full collected pool each time, rather than permanently pruning at the whole-task level — step-scoped filtering keeps later-step material recoverable.
When: A multi-step agent pipeline filters collected materials before each model call.
rag-over-tools-topk-injectionquery-shapeWhen an agent has many tools, embed the tool descriptions and inject only the top-K definitions semantically matched to the current task step — never all tool definitions at once.
When: An agent's tool inventory is large enough that injecting every definition wastes budget or degrades tool selection.
stable-prefix-cache-markingdisciplineKeep stable content (system instructions, project context) as a byte-identical prefix and mark it as a prompt-cache candidate (e.g., Anthropic cache_control) so repeated calls reuse the cache instead of re-paying for the prefix.
When: A pipeline makes repeated model calls sharing the same leading instructions/context on an API with prompt caching.
three-layer-memory-injectioncompositeCollect agent memory in three layers with distinct injection rules: short-term (recent turns verbatim up to a threshold, then summarized), working (a small scratchpad JSON of plan status/current step/intermediate results/error state, always injected), long-term (persistent store queried semantically, top-K injected only above a similarity threshold).
When: Designing the memory-gathering half of an agent's per-call context collection.
tiered-history-compression-laddercalibrationCompress conversation history on a three-level ladder triggered by a token threshold (e.g., 4K): keep the last K turns verbatim; summarize earlier turns into bullet-point key decisions; aggressively reduce to key facts, decisions made, and current state.
When: Accumulated conversation history is approaching or past the pipeline's history token threshold.
when-in-doubt-keep-at-filtercalibrationIn a re-runnable context-filtering stage, bias toward keeping doubtful items: filtering errors are recoverable by re-running, but context missing from the call is not.
When: A filter stage must decide whether a borderline item is relevant enough to enter the model's context.
Get the whole skill
All 8 primitives of this skill as one package, with the order to apply them.
Buy only the primitives you need
Each primitive is 1 credit (≈ €0.10). Pick them from the list above — the button is next to each one.
Upgrade your own skill
Paste your skill; we pick the 5 primitives from the shelf that fit it best, as one bundle for 5 credits (≈ €0.50).
Upgrade my skillNeighbour skills
Skills whose primitives are closest to this one (bge-m3 similarity):