Any company that answers customers at volume keeps a rulebook for what its agents may write. Refunds have thresholds and some promises are off limits. A general-purpose model sees that rulebook only if the whole policy rides along in every prompt, on every call.
This page trains the policy into the model itself. A small model learns it live in your browser tab, through the same loop that Thinking Machines' Tinker service runs on its open-weights Inkling-Small. Larkfield, the retailer in the examples, is fictional.
Each square is one reply the model wrote, scored the moment it was written. Each column is one training step, read left to right.
Paste a reply to see which checks it fails and the reward the training loop would give it. The same scorer runs the live training and the Tinker kit.
Paste past replies to see which pass. The compliant ones become a training file, so a first supervised round can teach the model your team's own wording.
The kit runs this loop on Thinking Machines' open-weights model from a laptop, while Tinker supplies the GPUs. You keep the trained adapter.
python evaluate.pyMeasure base Inkling-Small on fresh cases, with and without the policy in its prompt, before any training.
python train.pyFine-tune a LoRA adapter with the policy as the reward. A default run comes to an estimated $3.56 at today's prices.
python evaluate.py --tunedCompare all four arms on trained kinds of case and on one kind training never saw.
Change a switch, then press Start over.
A three-layer transformer with weights, pretrained on synthetic replies for Larkfield, a fictional retailer. It makes the right refund call about seven times in ten because it has never seen the policy.
Each step scores 16 sampled replies against the policy and the case facts. The scores update a rank-8 LoRA adapter and a bias on the output layer, as train.py does on Tinker. Plan renewal stays out of training as a held-out test.
Fully compliant replies on the case types the adapter trained on. Right refund decisions went from 70% to 95%. Each run passed 80% compliant, as a ten-step average, by step 57.
Wording failures fell from the first step. Refund decisions sat near 70% for about 30 steps, then climbed.
Plan renewals stayed out of training. Right decisions on them fell from 75% to 67%, since nothing taught the adapter their 14-day window.
With the decision check off, the training reward reached 0.98. Refund decisions stayed at the base model's 70%.
Without the base-model penalty, the base model's score for the replies fell from −0.18 to −0.48 per word over 160 steps.
| Check | Kind | Penalty | What it says |
|---|
Reward is exp(−penalty ÷ 4), scaled down outside 45 to 110 words. A reply with no penalty inside the word band counts as fully compliant.
Put a line of three hyphens between replies. Nothing pasted here leaves this browser.
| Reply | Words | Verdict | Failed checks |
|---|
The replies that pass, in the chat format Tinker's supervised recipe reads. The kit's transcripts.py builds the same file from a folder. Case facts are unknown here, so the refund decision goes unchecked.
This demonstration policy is 374 Inkling tokens, about 22 cents per thousand replies at $0.58 per million prompt tokens. A real policy manual runs far longer, and the prompt carries all of it on every call.
An estimate for 40 steps of 128 replies at 140 tokens a reply, thinking effort 0, using Inkling-Small's discounted Tinker prices per million tokens on 6 October 2026. List prices are double.
sampler = training_client.save_weights_and_get_sampling_client()
for case in batch:
replies = sampler.sample(case.prompt, num_samples=8, ...)
rewards = [policy.reward(text(r), case.eligible) for r in replies]
mean = sum(rewards) / len(rewards)
datums += [make_datum(case.prompt, r, x - mean)
for r, x in zip(replies, rewards)]
training_client.forward_backward(datums, "importance_sampling")
training_client.optim_step(AdamParams(learning_rate=4e-5))
Abridged from train.py in the kit.
Each case gives the model what an agent sees on screen and the customer's message. The policy thresholds stay inside the reward, so a trained model has to find them by trial.
The evaluation runs every arm on fresh cases of the trained kinds and on one kind of case that training never saw.
Any message a company sends in volume under rules a machine can check fits this method. Payment reminders, dispute responses, claim decisions and supplier letters all qualify.
The policy file is the part you replace. Most checks are patterns with a penalty attached. The refund rule compares the decision in a reply with the facts of the case. Anything that resists being written as a check stays with a human reviewer or the optional grading model in the kit.
The small model's pretraining and seeded runs ran here, along with tests that match the browser engine to a numpy reference and the Python scorer to the JavaScript one. The Tinker kit ran against stand-in clients built on the real Tinker SDK types and the real Inkling tokenizer.
No run has touched the live Tinker service yet, so this page holds no Inkling weights or Inkling replies. The small model is about a millionth the size of Inkling-Small and learned from a corpus written so that a compliant wording always exists.