Most of the questions I ask about our services during a normal day are not hard. They are just scattered. What is this service's id? Which repo holds that one's code? What is the namespace for this thing in the Tokyo prod cluster? Is it sync or async? Each answer is sitting in some deploy repo, some values.yaml, some runbook, and I spend more time finding it than using it.

The service skills I have been building with coding agents (starting with the endpoint-finding skill) already know where all of this lives. So I tried one more thing: a small local model I can just ask, in plain English, with typos, and get the exact value back.

I built it with AI coding agents. I set the goal and made the calls along the way; the agents mapped the repos, wrote the extractors and the answer engine, generated the training data, ran the fine-tunes, and drafted the first version of this write-up and its diagrams. All service names, ids and hosts below are made up; the shapes, numbers and model results are real.

❯ service id for spot-fx service?
ml:spot-fix  [service_id]
Service id: svc-3b7e1d90a4c2458f9e6b2d70c18f5a44
  version id: ver-8c21f0e4b97d4a03a5e6c1d2f3087b19
  engine ref: eng-5d0a9e71c3b84f26b8a1e4c7d9026f3e
  engine definition id: engdef-a47c2e19f06b4d58931e7b2c8d5f0a64
  ↳ identity/service_id 87% | env any 94% | service ml:spot-fix 94%
TL;DR
  • Every value comes from a repo file. A small classifier (Laya) decides what you are asking; a fact index built from the service repos supplies the answer.
  • About 200 services on two internal platforms, about 600 deployments and 56 question types, from ids and repos to namespaces, logs, alerts and "how do I add an env".
  • Fine-tuning took full-question accuracy from 10% to 95% on 80 realistic test questions, and from 15% to 90% on held-out generated ones.
  • It reuses the repo cache the skills already maintain (~/.cache/service-endpoint) and re-clones a service's repos on demand when they are older than 24 hours.
  • Answers take about 0.3 s on a MacBook and about 41 ms on an A100.

01Why not just ask an LLM?

The things I ask about services are exact strings: a service id, a sixty-character hashed Kubernetes namespace, a log sourcetype, a gateway URL. A generative model that is right 95% of the time is wrong in exactly the way that hurts: it produces a plausible id, and I only find out when kubectl says Forbidden.

So the design splits the problem in two.

JobDone byCan it make something up?
Understand the question: which service, which fact, which environmentLaya, a non-generative classifierNo. It can only choose among the options it is given
Produce the valueA fact index extracted from the reposNo. It prints what the repo says

Laya is a "System 1" decision model. You give it a state (my question) and typed questions (choice, score, yes/no), and it returns calibrated probabilities over the options in a single forward pass. It never generates text. The English checkpoint is ModernBERT-large plus a small decision head, 421M parameters in all.

The worst Laya can do is pick the wrong option: the wrong field or the wrong service. Because its probabilities are calibrated, the system can see when that is likely and say "did you mean…" instead of guessing.

System at a glance
Offline: build the index, train the model ~/.cache/ service-endpoint deploy + app repos (shared with the skills) index_ml.py + index_web.py extractors data/index.json about 200 service records datagen.py synthetic questions models/laya base checkpoint finetune.py RLCD + calibration models/ laya-svc fine-tuned weights facts Online: one question, ~0.3 s question name shortlist + regex slots Laya 1 forward pass answer.py render the field answer
Laya only decides what you are asking; every value in the answer is read from the index built from the repos.

The rest of the post walks that picture left to right: getting the data, the JSON it becomes, what happens when you ask, generating training questions, fine-tuning, and the results.

02Getting the data

Start from what the skills already know

The endpoint skill keeps shallow clones of every deploy and app repo in ~/.cache/service-endpoint, refreshed through gh when they go stale. That cache became the raw material. svcask never makes a second copy.

Our services live on two platforms with very different layouts. On the ML platform, every service has its own deploy repo and app repo, spread over two GitHub orgs. On the web platform, every service is a Helm chart directory inside one deploy monorepo, with code in a separate monorepo. Before any extractor was written, the agents sent out three read-only scouts to map where every fact lives, with measured coverage: one over the ML-platform repos, one over the web platform's two monorepos, and one over every skill to list the questions each one can answer and the exact function or doc section that answers it. That catalog became both the extraction spec and the question list.

Where each fact comes from
ML platform service (two GitHub orgs) .deploy-info.json serviceId, engine ids service.json envs, regions, Sync/Worker .platform.yaml owner, rbac, cost centers, catalog id argo/values.yaml + Chart.yaml app repo, vault secrets, chart version charts/<svc>/<env>/<region>/values.yaml cluster, namespace, GPU, scaling, queues, Splunk templates/prometheus_rule.yaml alert rules app repo engine_definition.json, Makefile, Dockerfile, test scripts web platform service platform-deploy charts/<svc>/<env>/<region>/values.yaml app_key, url_paths, workers, init_script charts/job-scheduler templates host formulas, GPU hosts, appIdMap CODEOWNERS, service-mapping.json owners, ServiceNow id platform-code monorepo registry.yaml, build/apps/app_*.py index_ml.py about 60 services index_web.py about 150 services runbook titles as aliases data/index.json about 200 services · about 600 deployments · ~15 s build incident-responder runbook index
Every fact has one authoritative file; the extractors read them from the skills' shared repo cache.

The full build takes about 15 seconds from a warm cache.

Keeping it fresh

The index is only as good as the cache, so every answer checks the service it is about before answering.

How the index stays fresh
no no yes yes question about service X cached repo older than TTL (24 h)? cache changed after X was extracted? (another skill refreshed it) common.cache_clone() the resolver skill's helper: locked, atomic swap, keeps the stale clone if gh/VPN fails re-extract X only save index.json answer from index observed: "is auto mask async?" re-cloned auto-mask-deploy + auto-mask-app, answered in 13.3 s; next questions 0.2 s
The index reuses the cache the skills already maintain, and never copies it.

It showed up in testing. Asking "is auto mask async?" found auto-mask's repos past the cache TTL, re-cloned auto-mask-deploy and auto-mask-app into the shared cache, and answered in 13.3 s. The next questions came back in 0.2 s again.

03The JSON: one record per service

Each service becomes one record in data/index.json. The schema is a contract in svcask/schema.py: both extractors emit exactly these keys, with null, [] or {} when a fact is unknown. Here is a record trimmed to one deployment (values anonymised):

{
 "key": "ml:spot-fix",
 "platform": "ml",
 "name": "spot-fix",
 "aliases": ["spot-fix-deploy", "spot_fix_v1", "spot-fix-v1", "spot fix deploy"],
 "ids": {
  "service_id": "svc-3b7e1d90a4c2458f9e6b2d70c18f5a44",
  "engine_id": "Classification:spot-fix:svc-3b7e1d90a4c2458f9e6b2d70c18f5a44",
  "catalog_id": "400117"
 },
 "repos": {
  "deploy": {"org": "ml-org", "name": "spot-fix-deploy",
             "url": "https://github.com/ml-org/spot-fix-deploy"},
  "app":    {"org": "ml-org", "name": "spot-fix",
             "url": "https://github.com/ml-org/spot-fix"}
 },
 "owners": {"owner": "jdoe", "rbac_write": "team-ml-writers"},
 "deployments": [{
   "env": "prod1", "tier": "prod", "region": "us-east",
   "cluster": "k8s-prod-use1",
   "namespace": "ns-ml-spot-fix-deploy--4e1a7c2b--7d90f3aa",
   "mode": "sync", "run_mode": "predict",
   "gateway": "https://ml-gateway.example.com/v2/predict",
   "gpu": {"family": "nvidia-a10", "count": 1},
   "scaling": {"min": 1, "max": 3, "cpu_target": 250, "queue_target": null, "autoscaler": "cpu"},
   "logs": {"index": "ml-services", "sourcetype": "spot-fix-prod1-us-east"}
 }],
 "signing": {"enabled": true,
             "vault_stage": "secret/ml/signing/spot-fix",
             "vault_prod": "secret/ml/signing-prod/spot-fix"},
 "alerts": [{"name": "SpotFixTooMany500s", "severity": "sev3", "for": "15m"}],
 "engine": {"inputs": ["image", "masks", "mask_offsets"], "params": ["sign_output"]}
}

The full record also carries sources (repo paths, mtimes and extraction time, which drive the freshness check), versions, routes, test and warnings. Warnings are data conflicts the extractor noticed, for example "cost centers in .platform.yaml disagree with service.json", and an answer only shows the ones about the field you asked for.

Deployments are per env × region × cluster, because that is how reality is shaped: one service's prod runs on two clusters with different platform versions. Every deployment-level answer can be narrowed by prod, stage1, ap-tokyo or k8s-stage-use1.

04What can you ask?

The agents turned the skill catalog into 9 categories and 56 attributes, each with a short description. The descriptions are exactly the options Laya sees, so they are kept terse: the option budget is 192 tokens per question.

CategoryAttributes (examples)
identityservice_id, engine_id, catalog_id, platform, name, description, owner, cost_center
codeapp_repo, deploy_repo, versions, image
endpointendpoint, public_endpoint, routes, health, debug_ui, curl
deploymentenvironments, cluster, namespace, mode (sync/async), gpu, scaling, queues
observabilitylogs, grafana, monitoring_hosts, alerts, argocd, newrelic, runbook
authtoken, vault, signing, kube_access
interfaceinputs, outputs, params
listinglist_all, by_region, by_mode, by_route, by_cluster, by_signing, by_owner
howtoadd env, alerts, test on pod, trace a request, signing, upload, review a PR, incident, vault, kubectl

On top of the attribute, Laya also picks the environment (any, dev, stage, prod or prerelease) and which service out of a shortlist of up to six candidates, or "none". Exact tokens are pulled out by regex, not the model: env cells like stage1, regions like ap-tokyo, clusters, routes like /v1/segment/pick-subject, and request ids. Exact strings are what regexes are for.

05What happens when you ask

One question, end to end
You nlu.py Laya fine-tuned answer.py index.json "service id for sopt-fix service?" regex slots (cells, regions, clusters, routes, request ids) → none shortlist by fuzzy alias match → [spot-fix 86.7] 12 choice questions, one forward pass (category, 9× attribute, env, service) identity 87% → service_id · env any 94% · spot-fix 94% Parse(attribute=service_id, service=ml:spot-fix, env=any) ensure fresh, then read ids.service_id svc-3b7e1d90… answer (+ "did you mean" if unsure)
The typo is absorbed by the fuzzy shortlist; Laya picks the field and the service; the value comes from the index.

The shortlist (no model). Every service name, alias, per-env app key and specific route segment is an alias. The question is scored against them with typo-tolerant matching (rapidfuzz, length-weighted, generic words like service or deploy count less), and the top six services above a threshold become Laya's options. Two lessons made this work:

  • Route segments are aliases only when they are specific. Every web-platform service exposes /endpoints, /docs and /status, so "endpoints" fuzzy-matched the word "endpoint" in every question about endpoints. Segments that are generic or shared by more than two services are dropped.
  • An alias cannot be another service's name. A copy-pasted template repo called its app repo auto-mask, which made "auto mask" ambiguous. Aliases that equal another service's canonical name are dropped.

One Laya forward pass. All questions for one query go into a single predict call: the category, the attribute question of every category, the environment and the service. The parser picks the best category first and then its best attribute. Training only ever asks the attribute question of the correct category, so the other categories' answers are noise and are never used on their own.

Guardrails. A low-confidence service pick lists the runner-ups ("did you mean …"). If Laya says "none" but one candidate is an unambiguous name or route match (score 90 or more and 15 points ahead of the next), svcask answers for it and says so. That is how "public endpoint for pick-subject service" resolves to image-core: pick-subject is a route of image-core, not a service. An uncertain attribute (under 30%) starts the answer with the alternative readings.

06Four questions, traced

Diagrams only go so far, so here are four questions run through the fine-tuned model, with the exact options Laya was given and the probabilities it returned. The probabilities are printed from the running system; only the names are swapped.

1. "service id for spot-fix service?"

One question, twelve choices
fixed text in schema.py (retrain to change) rebuilt from index.json on every question bar length = probability (full track = 1.0) “service id for spot-fix service?” Laya · one forward pass the other 8 attr_* questions are near-uniform (max 0.12–0.34): never trained, so ignored category· 9 fixed options schema.py identity 0.934 code 0.008 endpoint 0.008 deployment 0.008 +5 more at 0.008 attr_identity· 8 fixed options schema.py service_id 0.935 engine_id 0.009 catalog_id 0.009 +5 more at 0.009 env· 5 fixed options schema.py any 0.936 prod 0.016 stage 0.016 dev 0.016 prerelease 0.016 service· built per query from the shortlist index.json ml:spot-fix 0.939 spot-fix (ml) aka spot-fix-deploy, spot_fix_v1 none 0.061 pick = P(identity) × P(service_id | identity) = 0.934 × 0.935 = 87% answer.py reads ids.service_id svc-3b7e1d90…
Laya only ever sees labelled options; three option sets are fixed, the service options come from the index at question time.

Laya gets twelve multiple-choice questions in one forward pass. Each option is a label plus a short description, and Laya returns a probability for every option:

QuestionOptionsWhat Laya returned
category9, fixedidentity 0.934, every other category 0.008
attr_identity8, fixedservice_id 0.935, engine_id 0.009, catalog_id 0.009, …
the other 8 attr_* questions3 to 10 each, fixednear-uniform (the largest single value is 0.34)
env5, fixedany 0.936, prod / stage / dev / prerelease 0.016 each
serviceshortlist + "none"ml:spot-fix 0.939, none 0.061

The service question is the only one whose options change per question. For this one the shortlist found a single candidate (lexical score 101), so Laya saw exactly this:

"service": {
  "type": "choice",
  "instructions": "Which service is the question about?",
  "criteria": {
    "ml:spot-fix": "spot-fix (ml) aka spot-fix-deploy, spot_fix_v1",
    "none": "no specific service, or none of these"
  }
}

The eight attribute questions for the other categories come out close to flat. That is expected: training only ever asks the attribute question of the correct category, so the model has no opinion elsewhere, which is exactly why the parser picks the category first. The final pick is P(identity) × P(service_id | identity) = 0.934 × 0.935 ≈ 87%, the number in the explain line at the top of this post. Then answer.py reads ids.service_id from the record.

2. "public endpoint for pick-subject service"

pick-subject is not a service. It is a route (/v1/segment/pick-subject) that the web-platform service image-core exposes, so the shortlist only finds it through a route alias:

StepResult
shortlistweb:image-core 91.8 (route alias "pick subject"), web:vector-tools 61.6
category / attributeendpoint 0.934 → public_endpoint 0.937
envany 0.936
servicenone 0.937, image-core 0.032, vector-tools 0.032

Laya is right to be unsure: nothing in the question names a service. This is where the guardrail earns its keep. image-core scored 91.8, above the 90 threshold and 30 points ahead of the runner-up, so svcask answers for it and says so:

(picked web:image-core by name/route match; Laya was unsure which service you meant)
web:image-core  [public_endpoint]
Public hosts:
  prod/us-west/k8s-prod-usw1: image-core.example.com, image-core-cdn.example.com
  …

3. "namespace of spot-fix in ap-tokyo prod"

Here the regex and the model split the work. The regex pulls out ap-tokyo as an exact region; Laya decides the tier.

StepResult
regex slotsregions = [ap-tokyo]
category / attributedeployment 0.934 → namespace 0.936
envprod 0.936
serviceml:spot-fix 0.939

The service has six deployments. The answer engine keeps the ones where tier = prod and region = ap-tokyo, which leaves exactly one, and prints it with a ready-to-run command:

ml:spot-fix  [namespace · prod]
Namespace / label:
  prod1/ap-tokyo/k8s-prod-apt1: ns-ml-spot-fix-deploy--b52c19e0--3fa8d614 -l gitrepo=ml-org--spot-fix
  kubectl --context k8s-prod-apt1 -n ns-ml-spot-fix-deploy--b52c19e0--3fa8d614 get pods -l gitrepo=ml-org--spot-fix

4. "which services run in ap-tokyo"

No service is named, so the shortlist is empty and Laya is not even asked the service question. It answers listing 0.934 → by_region 0.936, and the listing runs over the whole index with the ap-tokyo slot as the filter:

15 match(es):
  ml:spot-fix: prod1, stage1
  ml:auto-mask: prod3, stage3, stage4
  ml:clutter-detect: prod2, stage2
  …
  ml:embed-gen: prod1

07Teaching Laya our domain

Out of the box, Laya is a generalist. Zero-shot, with 56 domain-specific attributes, it got 10% of the realistic questions fully right. That was a surprise, because the very first sanity check used four options and looked perfect. Always test with the real option set.

Fine-tuning needs labelled examples, so the agents generated them from the index itself.

How training data is made
train held-out index.json about 200 services 828 phrasing templates (≥14 per attribute) most train services 660 train templates 168 held-out templates ~20 held-out services fill slots casual names, typos, env / cell / region words, Slack-style filler shortlist exactly like inference (25% candidate order shuffled) train.jsonl 40,000 rows → 150,167 items soft target: 0.94 on the correct option fill slots same as the train side shortlist same as the train side eval.jsonl never-seen templates AND services
Evaluation uses phrasings and services the model never saw, so the accuracy numbers are honest.

Three decisions mattered:

  • Honest evaluation. The eval set uses templates and services the model never saw. Training shortlists are built only from train services, so a held-out service never appears even as a distractor. There is also an independent set of 80 realistic questions, written separately from the training templates and in a different style, including the three that started all this.
  • Train on inference-shaped inputs. Service candidates come from the same shortlist function the live system uses, so if the right service is not in the shortlist, the label is "none", just like at runtime. A quarter of the rows shuffle the candidate order; otherwise the model learns "always pick the first one".
  • Soft targets. Laya's training reads a distribution per question: 0.94 on the correct option and the rest spread evenly.

A training row, trimmed (names swapped):

{
 "state": "incident response steps for job-sceduler-preview on stage thanks!",
 "questions": {
  "category": {"type": "choice",
               "instructions": "What kind of information is the question asking for?",
               "criteria": {"identity": "ids, name, platform, owner or description of a service",
                            "code": "source repo, deploy repo, versions or docker image",
                            "…": "7 more"}},
  "service": {"type": "choice",
              "instructions": "Which service is the question about?",
              "criteria": {"web:job-scheduler-preview": "job-scheduler-preview (web) aka job_scheduler, …",
                           "web:job-scheduler": "job-scheduler (web) aka …",
                           "web:latency-probe": "…", "web:text-gen-partner": "…",
                           "none": "no specific service, or none of these"}}
 },
 "gold": {
  "category":   {"probabilities": {"howto": 0.94, "...": "rest spread"}},
  "attr_howto": {"probabilities": {"howto_incident": 0.94}},
  "env":        {"probabilities": {"stage": 0.94}},
  "service":    {"probabilities": {"web:job-scheduler-preview": 0.94}}
 },
 "meta": {"template_id": "howto_incident#7", "split": "train", "n_candidates": 4}
}

Note the typo ("sceduler") and the near-duplicate distractor (job-scheduler). That is exactly what the shortlist produces in real use.

08Fine-tuning

An agent adapted Laya's own single-device fine-tuning script (Apache-2.0). Its method, RLCD, combines a policy-gradient term rewarded by proper scoring rules with a soft cross-entropy against the target distribution, then fits calibration temperatures at the end.

The fine-tuning loop
next step every epoch after the last epoch batch of items query + one typed question + soft target forward pass: logits per option ModernBERT-large encoder + decision head, 421M policy-gradient term 4 noisy logit samples, reward = proper scoring rules (spherical 0.75 + ranked-prob 1.0) soft cross-entropy vs the target distribution loss loss goes negative: it includes a reward term AdamW encoder 2.5e-5 · head 1e-4 · cosine · grad clip 1.0 checkpoint_latest fit one temperature per question type on a held-out slice models/laya-svc
Proper scoring rules reward honest probabilities, which is why answer confidences mean something.

Proper scoring rules reward honest probabilities, which is why the confidence numbers in answers mean something: about 0.93 when the model is right, lower when it is wrong. The calibration slice is held out by row, so the same question never shows up in both training and calibration. And yes, the loss goes negative during training. It includes a reward term, so falling means improving.

The fine-tuned model was trained on a single rented A100: 40,000 generated questions (150,167 training items), 6 epochs at batch size 32 in bf16, about an hour and a half end to end.

09Results

"Fully right" means the attribute, environment and service are all correct, which is what you need for a correct answer.

Accuracy, base vs fine-tuned
base Laya fine-tuned Laya 0% 25% 50% 75% 100% Realistic questions (80), fully right 10.1% 94.9% Held-out generated questions (3,000), fully right 14.9% 89.5% Service picked correctly (realistic) 20.3% 98.4%
“Fully right” means attribute, environment and service are all correct; held-out = templates and services the model never trained on.

The 80 realistic questions are independent of the training templates:

ModelFully rightCategoryAttributeEnvService
Base Laya (zero-shot)10.1%43.0%35.4%88.6%20.3%
Fine-tuned Laya94.9%98.7%97.5%98.7%98.4%

On 3,000 held-out generated questions (phrasings and services it never trained on), the fine-tuned model is fully right 89.5% of the time against 14.9% for the base model, and picks the right service 99.8% of the time against 43.2%.

The fine-tuned model's four remaining misses on the realistic set:

QuestionExpectedGot
public endpoint for pick-subject servicepublic_endpoint · image-corepublic_endpoint · no service; the name/route fallback still answers image-core
where does the code for auto mask liveapp_repo · any envapp_repo · prod: the right answer, with an extra env filter
object erase prod instance typesgpucluster
do we page on auto-mask prod errorsalertshealth

These numbers do not include the name/route fallback, so the live system does a little better. The misses are neighbouring intents that share words ("instance" and cluster, "page on errors" and health); more phrasing templates for those pairs is the obvious next step.

10A new service ships: do I retrain?

No. A new service only needs the index rebuilt. Retraining is for new kinds of questions.

A new service ships
new deploy repo / web app merged python -m svcask.index_build --refresh (or :rebuild in the REPL) · ~15 s discovery rules ML: deploy repo has .deploy-info.json + service.json (re-lists the two ML-platform orgs via gh) web: a directory under charts/ in platform-deploy (pulled fresh) index.json gains the record next question: the shortlist offers it as an option Laya picks among options (held-out services: right 2,029 / 2,031) answer per-answer refresh updates known services only; it never discovers new ones nickname appears in no repo add it as an alias no retraining new attribute / category / reworded option in schema.py datagen → finetune → evaluate the only case that retrains
Adding a service is an index rebuild, not a training run.

Why no retraining is needed:

  • Laya never learned a fixed list of services. For each question the shortlist pulls up to six candidates from data/index.json, and those candidates are the options Laya chooses from (the service block in example 1). A service that is in the index shows up as an option automatically.
  • Answers come from the index, not the model, so the new service's ids, endpoints and repos appear the moment the index has them.

What you do need is to let the index discover the new repo:

.venv/bin/python -m svcask.index_build --refresh      # ~15 s warm, or :rebuild inside the REPL

--refresh re-lists the ML platform's two orgs through gh and pulls the web platform's deploy monorepo fresh. A new ML-platform service is picked up once its deploy repo has .deploy-info.json and service.json; a new web-platform service once it has a directory under charts/ in the deploy monorepo. The automatic per-answer refresh only updates services already in the index; it never discovers new ones.

You would retrain in two cases:

  • The question taxonomy changes: adding, renaming or rewording an attribute or category in svcask/schema.py. Those option texts are exactly what Laya was trained on.
  • You want a new kind of question answered. A new renderer in answer.py is not enough on its own; the model needs training examples that ask for it.

And a nickname the matcher cannot find, say a product name that appears in no repo, is not a training problem either. Add it as an alias and it becomes findable without touching the model.

What I like most is where it sits: the model only has to understand me, and the repos stay the source of truth. When a deploy repo changes, the answer changes with it, and nobody has to retrain anything.