Most of the questions I ask about our services during a normal day are not hard. They are just scattered. What is this service's id? Which repo holds that one's code? What is the namespace for this thing in the Tokyo prod cluster? Is it sync or async? Each answer is sitting in some deploy repo, some values.yaml, some runbook, and I spend more time finding it than using it.
The service skills I have been building with coding agents (starting with the endpoint-finding skill) already know where all of this lives. So I tried one more thing: a small local model I can just ask, in plain English, with typos, and get the exact value back.
I built it with AI coding agents. I set the goal and made the calls along the way; the agents mapped the repos, wrote the extractors and the answer engine, generated the training data, ran the fine-tunes, and drafted the first version of this write-up and its diagrams. All service names, ids and hosts below are made up; the shapes, numbers and model results are real.
❯ service id for spot-fx service?
ml:spot-fix [service_id]
Service id: svc-3b7e1d90a4c2458f9e6b2d70c18f5a44
version id: ver-8c21f0e4b97d4a03a5e6c1d2f3087b19
engine ref: eng-5d0a9e71c3b84f26b8a1e4c7d9026f3e
engine definition id: engdef-a47c2e19f06b4d58931e7b2c8d5f0a64
↳ identity/service_id 87% | env any 94% | service ml:spot-fix 94%
- Every value comes from a repo file. A small classifier (Laya) decides what you are asking; a fact index built from the service repos supplies the answer.
- About 200 services on two internal platforms, about 600 deployments and 56 question types, from ids and repos to namespaces, logs, alerts and "how do I add an env".
- Fine-tuning took full-question accuracy from 10% to 95% on 80 realistic test questions, and from 15% to 90% on held-out generated ones.
- It reuses the repo cache the skills already maintain (
~/.cache/service-endpoint) and re-clones a service's repos on demand when they are older than 24 hours. - Answers take about 0.3 s on a MacBook and about 41 ms on an A100.
01Why not just ask an LLM?
The things I ask about services are exact strings: a service id, a sixty-character hashed Kubernetes namespace, a log sourcetype, a gateway URL. A generative model that is right 95% of the time is wrong in exactly the way that hurts: it produces a plausible id, and I only find out when kubectl says Forbidden.
So the design splits the problem in two.
| Job | Done by | Can it make something up? |
|---|---|---|
| Understand the question: which service, which fact, which environment | Laya, a non-generative classifier | No. It can only choose among the options it is given |
| Produce the value | A fact index extracted from the repos | No. It prints what the repo says |
Laya is a "System 1" decision model. You give it a state (my question) and typed questions (choice, score, yes/no), and it returns calibrated probabilities over the options in a single forward pass. It never generates text. The English checkpoint is ModernBERT-large plus a small decision head, 421M parameters in all.
The worst Laya can do is pick the wrong option: the wrong field or the wrong service. Because its probabilities are calibrated, the system can see when that is likely and say "did you mean…" instead of guessing.
The rest of the post walks that picture left to right: getting the data, the JSON it becomes, what happens when you ask, generating training questions, fine-tuning, and the results.
02Getting the data
Start from what the skills already know
The endpoint skill keeps shallow clones of every deploy and app repo in ~/.cache/service-endpoint, refreshed through gh when they go stale. That cache became the raw material. svcask never makes a second copy.
Our services live on two platforms with very different layouts. On the ML platform, every service has its own deploy repo and app repo, spread over two GitHub orgs. On the web platform, every service is a Helm chart directory inside one deploy monorepo, with code in a separate monorepo. Before any extractor was written, the agents sent out three read-only scouts to map where every fact lives, with measured coverage: one over the ML-platform repos, one over the web platform's two monorepos, and one over every skill to list the questions each one can answer and the exact function or doc section that answers it. That catalog became both the extraction spec and the question list.
The full build takes about 15 seconds from a warm cache.
Keeping it fresh
The index is only as good as the cache, so every answer checks the service it is about before answering.
It showed up in testing. Asking "is auto mask async?" found auto-mask's repos past the cache TTL, re-cloned auto-mask-deploy and auto-mask-app into the shared cache, and answered in 13.3 s. The next questions came back in 0.2 s again.
03The JSON: one record per service
Each service becomes one record in data/index.json. The schema is a contract in svcask/schema.py: both extractors emit exactly these keys, with null, [] or {} when a fact is unknown. Here is a record trimmed to one deployment (values anonymised):
{
"key": "ml:spot-fix",
"platform": "ml",
"name": "spot-fix",
"aliases": ["spot-fix-deploy", "spot_fix_v1", "spot-fix-v1", "spot fix deploy"],
"ids": {
"service_id": "svc-3b7e1d90a4c2458f9e6b2d70c18f5a44",
"engine_id": "Classification:spot-fix:svc-3b7e1d90a4c2458f9e6b2d70c18f5a44",
"catalog_id": "400117"
},
"repos": {
"deploy": {"org": "ml-org", "name": "spot-fix-deploy",
"url": "https://github.com/ml-org/spot-fix-deploy"},
"app": {"org": "ml-org", "name": "spot-fix",
"url": "https://github.com/ml-org/spot-fix"}
},
"owners": {"owner": "jdoe", "rbac_write": "team-ml-writers"},
"deployments": [{
"env": "prod1", "tier": "prod", "region": "us-east",
"cluster": "k8s-prod-use1",
"namespace": "ns-ml-spot-fix-deploy--4e1a7c2b--7d90f3aa",
"mode": "sync", "run_mode": "predict",
"gateway": "https://ml-gateway.example.com/v2/predict",
"gpu": {"family": "nvidia-a10", "count": 1},
"scaling": {"min": 1, "max": 3, "cpu_target": 250, "queue_target": null, "autoscaler": "cpu"},
"logs": {"index": "ml-services", "sourcetype": "spot-fix-prod1-us-east"}
}],
"signing": {"enabled": true,
"vault_stage": "secret/ml/signing/spot-fix",
"vault_prod": "secret/ml/signing-prod/spot-fix"},
"alerts": [{"name": "SpotFixTooMany500s", "severity": "sev3", "for": "15m"}],
"engine": {"inputs": ["image", "masks", "mask_offsets"], "params": ["sign_output"]}
}
The full record also carries sources (repo paths, mtimes and extraction time, which drive the freshness check), versions, routes, test and warnings. Warnings are data conflicts the extractor noticed, for example "cost centers in .platform.yaml disagree with service.json", and an answer only shows the ones about the field you asked for.
Deployments are per env × region × cluster, because that is how reality is shaped: one service's prod runs on two clusters with different platform versions. Every deployment-level answer can be narrowed by prod, stage1, ap-tokyo or k8s-stage-use1.
04What can you ask?
The agents turned the skill catalog into 9 categories and 56 attributes, each with a short description. The descriptions are exactly the options Laya sees, so they are kept terse: the option budget is 192 tokens per question.
| Category | Attributes (examples) |
|---|---|
| identity | service_id, engine_id, catalog_id, platform, name, description, owner, cost_center |
| code | app_repo, deploy_repo, versions, image |
| endpoint | endpoint, public_endpoint, routes, health, debug_ui, curl |
| deployment | environments, cluster, namespace, mode (sync/async), gpu, scaling, queues |
| observability | logs, grafana, monitoring_hosts, alerts, argocd, newrelic, runbook |
| auth | token, vault, signing, kube_access |
| interface | inputs, outputs, params |
| listing | list_all, by_region, by_mode, by_route, by_cluster, by_signing, by_owner |
| howto | add env, alerts, test on pod, trace a request, signing, upload, review a PR, incident, vault, kubectl |
On top of the attribute, Laya also picks the environment (any, dev, stage, prod or prerelease) and which service out of a shortlist of up to six candidates, or "none". Exact tokens are pulled out by regex, not the model: env cells like stage1, regions like ap-tokyo, clusters, routes like /v1/segment/pick-subject, and request ids. Exact strings are what regexes are for.
05What happens when you ask
The shortlist (no model). Every service name, alias, per-env app key and specific route segment is an alias. The question is scored against them with typo-tolerant matching (rapidfuzz, length-weighted, generic words like service or deploy count less), and the top six services above a threshold become Laya's options. Two lessons made this work:
- Route segments are aliases only when they are specific. Every web-platform service exposes
/endpoints,/docsand/status, so "endpoints" fuzzy-matched the word "endpoint" in every question about endpoints. Segments that are generic or shared by more than two services are dropped. - An alias cannot be another service's name. A copy-pasted template repo called its app repo
auto-mask, which made "auto mask" ambiguous. Aliases that equal another service's canonical name are dropped.
One Laya forward pass. All questions for one query go into a single predict call: the category, the attribute question of every category, the environment and the service. The parser picks the best category first and then its best attribute. Training only ever asks the attribute question of the correct category, so the other categories' answers are noise and are never used on their own.
Guardrails. A low-confidence service pick lists the runner-ups ("did you mean …"). If Laya says "none" but one candidate is an unambiguous name or route match (score 90 or more and 15 points ahead of the next), svcask answers for it and says so. That is how "public endpoint for pick-subject service" resolves to image-core: pick-subject is a route of image-core, not a service. An uncertain attribute (under 30%) starts the answer with the alternative readings.
06Four questions, traced
Diagrams only go so far, so here are four questions run through the fine-tuned model, with the exact options Laya was given and the probabilities it returned. The probabilities are printed from the running system; only the names are swapped.
1. "service id for spot-fix service?"
Laya gets twelve multiple-choice questions in one forward pass. Each option is a label plus a short description, and Laya returns a probability for every option:
| Question | Options | What Laya returned |
|---|---|---|
| category | 9, fixed | identity 0.934, every other category 0.008 |
| attr_identity | 8, fixed | service_id 0.935, engine_id 0.009, catalog_id 0.009, … |
| the other 8 attr_* questions | 3 to 10 each, fixed | near-uniform (the largest single value is 0.34) |
| env | 5, fixed | any 0.936, prod / stage / dev / prerelease 0.016 each |
| service | shortlist + "none" | ml:spot-fix 0.939, none 0.061 |
The service question is the only one whose options change per question. For this one the shortlist found a single candidate (lexical score 101), so Laya saw exactly this:
"service": {
"type": "choice",
"instructions": "Which service is the question about?",
"criteria": {
"ml:spot-fix": "spot-fix (ml) aka spot-fix-deploy, spot_fix_v1",
"none": "no specific service, or none of these"
}
}
The eight attribute questions for the other categories come out close to flat. That is expected: training only ever asks the attribute question of the correct category, so the model has no opinion elsewhere, which is exactly why the parser picks the category first. The final pick is P(identity) × P(service_id | identity) = 0.934 × 0.935 ≈ 87%, the number in the explain line at the top of this post. Then answer.py reads ids.service_id from the record.
2. "public endpoint for pick-subject service"
pick-subject is not a service. It is a route (/v1/segment/pick-subject) that the web-platform service image-core exposes, so the shortlist only finds it through a route alias:
| Step | Result |
|---|---|
| shortlist | web:image-core 91.8 (route alias "pick subject"), web:vector-tools 61.6 |
| category / attribute | endpoint 0.934 → public_endpoint 0.937 |
| env | any 0.936 |
| service | none 0.937, image-core 0.032, vector-tools 0.032 |
Laya is right to be unsure: nothing in the question names a service. This is where the guardrail earns its keep. image-core scored 91.8, above the 90 threshold and 30 points ahead of the runner-up, so svcask answers for it and says so:
(picked web:image-core by name/route match; Laya was unsure which service you meant)
web:image-core [public_endpoint]
Public hosts:
prod/us-west/k8s-prod-usw1: image-core.example.com, image-core-cdn.example.com
…
3. "namespace of spot-fix in ap-tokyo prod"
Here the regex and the model split the work. The regex pulls out ap-tokyo as an exact region; Laya decides the tier.
| Step | Result |
|---|---|
| regex slots | regions = [ap-tokyo] |
| category / attribute | deployment 0.934 → namespace 0.936 |
| env | prod 0.936 |
| service | ml:spot-fix 0.939 |
The service has six deployments. The answer engine keeps the ones where tier = prod and region = ap-tokyo, which leaves exactly one, and prints it with a ready-to-run command:
ml:spot-fix [namespace · prod] Namespace / label: prod1/ap-tokyo/k8s-prod-apt1: ns-ml-spot-fix-deploy--b52c19e0--3fa8d614 -l gitrepo=ml-org--spot-fix kubectl --context k8s-prod-apt1 -n ns-ml-spot-fix-deploy--b52c19e0--3fa8d614 get pods -l gitrepo=ml-org--spot-fix
4. "which services run in ap-tokyo"
No service is named, so the shortlist is empty and Laya is not even asked the service question. It answers listing 0.934 → by_region 0.936, and the listing runs over the whole index with the ap-tokyo slot as the filter:
15 match(es): ml:spot-fix: prod1, stage1 ml:auto-mask: prod3, stage3, stage4 ml:clutter-detect: prod2, stage2 … ml:embed-gen: prod1
07Teaching Laya our domain
Out of the box, Laya is a generalist. Zero-shot, with 56 domain-specific attributes, it got 10% of the realistic questions fully right. That was a surprise, because the very first sanity check used four options and looked perfect. Always test with the real option set.
Fine-tuning needs labelled examples, so the agents generated them from the index itself.
Three decisions mattered:
- Honest evaluation. The eval set uses templates and services the model never saw. Training shortlists are built only from train services, so a held-out service never appears even as a distractor. There is also an independent set of 80 realistic questions, written separately from the training templates and in a different style, including the three that started all this.
- Train on inference-shaped inputs. Service candidates come from the same shortlist function the live system uses, so if the right service is not in the shortlist, the label is "none", just like at runtime. A quarter of the rows shuffle the candidate order; otherwise the model learns "always pick the first one".
- Soft targets. Laya's training reads a distribution per question: 0.94 on the correct option and the rest spread evenly.
A training row, trimmed (names swapped):
{
"state": "incident response steps for job-sceduler-preview on stage thanks!",
"questions": {
"category": {"type": "choice",
"instructions": "What kind of information is the question asking for?",
"criteria": {"identity": "ids, name, platform, owner or description of a service",
"code": "source repo, deploy repo, versions or docker image",
"…": "7 more"}},
"service": {"type": "choice",
"instructions": "Which service is the question about?",
"criteria": {"web:job-scheduler-preview": "job-scheduler-preview (web) aka job_scheduler, …",
"web:job-scheduler": "job-scheduler (web) aka …",
"web:latency-probe": "…", "web:text-gen-partner": "…",
"none": "no specific service, or none of these"}}
},
"gold": {
"category": {"probabilities": {"howto": 0.94, "...": "rest spread"}},
"attr_howto": {"probabilities": {"howto_incident": 0.94}},
"env": {"probabilities": {"stage": 0.94}},
"service": {"probabilities": {"web:job-scheduler-preview": 0.94}}
},
"meta": {"template_id": "howto_incident#7", "split": "train", "n_candidates": 4}
}
Note the typo ("sceduler") and the near-duplicate distractor (job-scheduler). That is exactly what the shortlist produces in real use.
08Fine-tuning
An agent adapted Laya's own single-device fine-tuning script (Apache-2.0). Its method, RLCD, combines a policy-gradient term rewarded by proper scoring rules with a soft cross-entropy against the target distribution, then fits calibration temperatures at the end.
Proper scoring rules reward honest probabilities, which is why the confidence numbers in answers mean something: about 0.93 when the model is right, lower when it is wrong. The calibration slice is held out by row, so the same question never shows up in both training and calibration. And yes, the loss goes negative during training. It includes a reward term, so falling means improving.
The fine-tuned model was trained on a single rented A100: 40,000 generated questions (150,167 training items), 6 epochs at batch size 32 in bf16, about an hour and a half end to end.
09Results
"Fully right" means the attribute, environment and service are all correct, which is what you need for a correct answer.
The 80 realistic questions are independent of the training templates:
| Model | Fully right | Category | Attribute | Env | Service |
|---|---|---|---|---|---|
| Base Laya (zero-shot) | 10.1% | 43.0% | 35.4% | 88.6% | 20.3% |
| Fine-tuned Laya | 94.9% | 98.7% | 97.5% | 98.7% | 98.4% |
On 3,000 held-out generated questions (phrasings and services it never trained on), the fine-tuned model is fully right 89.5% of the time against 14.9% for the base model, and picks the right service 99.8% of the time against 43.2%.
The fine-tuned model's four remaining misses on the realistic set:
| Question | Expected | Got |
|---|---|---|
| public endpoint for pick-subject service | public_endpoint · image-core | public_endpoint · no service; the name/route fallback still answers image-core |
| where does the code for auto mask live | app_repo · any env | app_repo · prod: the right answer, with an extra env filter |
| object erase prod instance types | gpu | cluster |
| do we page on auto-mask prod errors | alerts | health |
These numbers do not include the name/route fallback, so the live system does a little better. The misses are neighbouring intents that share words ("instance" and cluster, "page on errors" and health); more phrasing templates for those pairs is the obvious next step.
10A new service ships: do I retrain?
No. A new service only needs the index rebuilt. Retraining is for new kinds of questions.
Why no retraining is needed:
- Laya never learned a fixed list of services. For each question the shortlist pulls up to six candidates from
data/index.json, and those candidates are the options Laya chooses from (theserviceblock in example 1). A service that is in the index shows up as an option automatically. - Answers come from the index, not the model, so the new service's ids, endpoints and repos appear the moment the index has them.
What you do need is to let the index discover the new repo:
.venv/bin/python -m svcask.index_build --refresh # ~15 s warm, or :rebuild inside the REPL
--refresh re-lists the ML platform's two orgs through gh and pulls the web platform's deploy monorepo fresh. A new ML-platform service is picked up once its deploy repo has .deploy-info.json and service.json; a new web-platform service once it has a directory under charts/ in the deploy monorepo. The automatic per-answer refresh only updates services already in the index; it never discovers new ones.
You would retrain in two cases:
- The question taxonomy changes: adding, renaming or rewording an attribute or category in
svcask/schema.py. Those option texts are exactly what Laya was trained on. - You want a new kind of question answered. A new renderer in
answer.pyis not enough on its own; the model needs training examples that ask for it.
And a nickname the matcher cannot find, say a product name that appears in no repo, is not a training problem either. Add it as an alias and it becomes findable without touching the model.
What I like most is where it sits: the model only has to understand me, and the repos stay the source of truth. When a deploy repo changes, the answer changes with it, and nobody has to retrain anything.
Comments