← the whole session plugin/skills/agentic-find-how-to-optimize-monthly-maintenence-only/SKILL.md
Tune a semantic institutional-search tool's signal quality during monthly maintenance, using agent-find as a worked example. Use when results feel noisy or need recalibration — but only once you have a comparable search tool of your own; this skill has nothing to tune if you don't.
Tuning a Semantic Institutional-Search Tool — Worked Example: agent-find
Before you use this
This entire skill is a worked example of tuning ONE specific, private search tool — agent-find (an institutional-memory search over your own sessions, docs, and commits; see /alignment-harness:harness-setup if you want to build something similar). If you haven't built or connected a comparable semantic-search / RAG tool over your own project's history, none of the file paths or tuning knobs below apply to you, and there's nothing here to do — skip this skill entirely.
If you have built your own equivalent (any tool that embeds and ranks your project's docs, sessions, commits, etc.), the method below transfers directly: source weighting, section thresholds, a fixed set of test queries, and a checklist for catching regressions. Replace every file path below with wherever your own tool's source lives.
Overview
agent-find is the canonical first search for all agents, in this reference setup. Its output quality directly determines whether agents act correctly on the first pass. This skill covers monthly tuning of signal/noise ratio, section thresholds, and source weights — as an illustration of the general practice, using one implementation's actual file layout.
Not for feature work. For adding new sections or capabilities, read your own tool's source directly.
Key Files (a reference layout — a worked example; point at your own tool's files instead)
| File | Purpose |
|---|---|
agent-swarm-mcp/mcp-server/agent-find.js |
Main search engine + output formatter |
agent-swarm-mcp/mcp-server/search-profiles.js |
Source weights per profile |
agent-swarm-mcp/scripts/agent-find-cli.js |
CLI test harness |
agent-swarm-mcp/logs/agent-artifacts.jsonl |
Session artifact JSONL (533+ entries) |
agent-swarm-mcp/data/vector-cache/ |
Embedding cache |
How to Test
Always test with the CLI — MCP transport may be down and results are identical:
cd /tmp && node <project>/agent-swarm-mcp/scripts/agent-find-cli.js "your query" 2>&1
Test with 3 different query types each maintenance cycle:
- An operational task:
"how do I add credit to a customer stripe account" - A debugging query:
"why is the stripe webhook not firing for new subscriptions" - A domain knowledge query:
"how to get trial conversion rates by offer"
Output Structure (What Good Looks Like)
Found N results for "..." (from X candidates):
Hits: intent(N) | artifact(N) | tool(N) | commit(N) | doc(N)
1-4. Full intent results with body snippet
─── MORE MATCHES ─── (max 5, intents only, one line each)
─── SESSION ARTIFACTS ─── (past work, agent_read ID, days ago)
─── COMMITS ─── (recent code changes in domain)
─── DOCS ─── (reference material)
─── SWARM ─── (pattern memories, deduped)
─── STRATEGY ─── (company intents, max 3, high threshold)
─── TOOLS MATCHED ─── (endpoints + skills, slug + %, deduped)
Signs it's working correctly:
- Top 4 results are specific UX intents for the domain
stripe-add-customer-creditor equivalent skill appears first in TOOLS MATCHED for operational queries- SESSION ARTIFACTS shows research reports with clean titles + IDs
- Company intents NOT in main results (only in STRATEGY section)
- No duplicate lines anywhere
Source Weights (search-profiles.js → PROFILES.default)
Current tuned values:
intent: 1.4 // highest — UX intents are ground truth
artifact: 1.0 // session work — high value
tool: 1.1 // endpoints/skills — operational
api: 1.1 // swagger endpoints
doc: 1.0 // reference material
swarm: 0.9 // pattern memories
commit: 0.9 // recent code changes
company_intent: 0.3 // strategic context ONLY — appears via STRATEGY section, not main results
ux: 0.8 // UX assignments
task: 0.6 // mostly dead — don't increase
founder: 0.0 // never in default results
Rule: company_intent must stay ≤ 0.3 or it floods main results. It belongs only in the compact STRATEGY section.
Section Thresholds (agent-find.js → compactSectionDefs)
{ src: "artifact", limit: 10, threshold: 0.35 } // min relevance to show
{ src: "commit", limit: 5, threshold: 0 }
{ src: "doc", limit: 5, threshold: 0 }
{ src: "swarm", limit: 10, threshold: 0 } // filtered in formatter to ≥35%
{ src: "company_intent", limit: 3, threshold: 0.65 } // only show solid strategic matches
TOOLS MATCHED threshold: 0.30 composite, top 20, deduped by slug.
Section-Only Sources
These sources never appear in main results or MORE MATCHES — they have dedicated sections:
const SECTION_ONLY = new Set(["tool", "api", "skill", "artifact", "swarm"]);
If tools or artifacts start bleeding into MORE MATCHES, check this set.
Common Issues & Fixes
Problem: Company intents flooding main results
Symptom: Results 1-8 are all 🏢 company intent entries
Fix: Check company_intent weight in search-profiles.js — should be 0.3 not higher
Problem: SESSION ARTIFACTS section empty
Symptom: artifact(0) in hit counts, no artifact section
Check: Was buildIndex() called on the vectorStore before agentFind()?
In scripts/agent-find-cli.js:
const vectorStore = new MemoryVectorStore({ logsDir: LOGS_DIR, cacheDir: VECTOR_CACHE_DIR });
await vectorStore.buildIndex(); // ← THIS MUST BE PRESENT
Problem: Artifact titles are raw session context
Symptom: Lines like "The user had three requests:" or "- Original request:"
Fix: Expand SKIP_RAW regex in formatAgentFindResults():
const SKIP_RAW = /^(The user|A user|User (is|was|had|request)|[-•]\s|\d+\.\s|Session (started|context)|...)/i;
Problem: TOOLS MATCHED has duplicates
Symptom: Same endpoint slug appears twice
Fix: Check seenTools Set dedup logic in formatter TOOLS MATCHED block
Problem: SWARM section shows founder knowledge (pricing/agency entries)
Symptom: "When considering agency partnerships..." appearing for coding queries
Fix: Swarm formatter filters item.more.includes("/founder") — verify this filter is active
Problem: Docs section has duplicate titles
Symptom: Same doc title appears twice with different paths
Fix: seenDoc Set dedup in formatter doc section — verify it's active
Problem: Swarm section appears with very low scores (15-25%)
Symptom: Swarm entries at 22% with no signal
Fix: Swarm minimum threshold in formatter: item.score >= 35
Artifact Detection Logic
Artifacts are session work auto-extracted by the artifact extractor. Detected in searchSwarm():
const isArtifact = m.agent_id === "artifact-extractor" ||
(m.tags || []).includes("auto-extracted");
Artifact more ref format: agent_read 10524 (5-digit ID)
Read full content: agent_read({ id: 10524 })
Monthly Tuning Checklist
- Run 3 test queries (operational, debugging, domain knowledge)
- Check hit count header — all expected sources appearing?
- Check top 4 results — are they specific intents for the domain or noise?
- Check TOOLS MATCHED — does the most relevant tool/endpoint appear first?
- Check SESSION ARTIFACTS — clean titles? relevant content? IDs present?
- Check for section bleed — tools/artifacts appearing in main results?
- Check for duplicates — in TOOLS MATCHED, DOCS, SWARM?
- Check STRATEGY section — appearing only when genuinely strategic? Max 3?
- If weights need adjustment — change
search-profiles.jsonly, test all 3 queries after - Commit changes — message format:
refine(agent-find): [what changed and why]
Architecture Notes
- Vector store searches JSONL logs only (
this.index) viasearch()method - Structured docs (intents, tools) indexed separately via
indexStructuredDocuments() - Both live in same
MemoryVectorStoreinstance but searched differently searchSwarm()is called withlimit * 4to ensure enough artifact candidates surface- Composite score =
relevance × sourceWeight × recency × statusMultiplier approvedstatus = 1.075x multiplier,proposed/pending= 0.925x