Rulebook
← dinand.com
A working demonstration of reinforcement fine-tuning

Teach a model your customer reply policy

Any company that answers customers at volume keeps a rulebook for what its agents may write. Refunds have thresholds and some promises are off limits. A general-purpose model sees that rulebook only if the whole policy rides along in every prompt, on every call.

This page trains the policy into the model itself. A small model learns it live in your browser tab, through the same loop that Thinking Machines' Tinker service runs on its open-weights Inkling-Small. Larkfield, the retailer in the examples, is fictional.

  1. PolicyNine banned habits, three required elements and a refund rule, each written as a check a machine can run.
  2. ScoreThe model drafts replies to customer cases. The checks turn each draft into a reward between 0 and 1.
  3. UpdateDrafts that beat their batch average pull a small adapter toward their wording.
  4. Repeat160 rounds of 16 drafts. The model's original weights never change.
Fully compliant replies
8%
Right refund decision
70%
Replies scored
0
step 0

Each square is one reply the model wrote, scored the moment it was written. Each column is one training step, read left to right.

Compliant Fails a wording or required check Wrong refund decision
Customer wrote
Policy rule, and where this case falls

Base model

Frozen weights, never shown the policy

Same model, with the adapter

Not trained yet

Compliance over the run

Fully compliantRight decision

Change a switch, then press Start over.

What the adapter changed

Words it now avoids
Words it now prefers

Failure rate by check

Base modelWith the adapterDecision checks

What runs here

A three-layer transformer with weights, pretrained on synthetic replies for Larkfield, a fictional retailer. It makes the right refund call about seven times in ten because it has never seen the policy.

Each step scores 16 sampled replies against the policy and the case facts. The scores update a rank-8 LoRA adapter and a bias on the output layer, as train.py does on Tinker. Plan renewal stays out of training as a held-out test.

Five seeded runs of 160 steps

8% → 93%

Fully compliant replies on the case types the adapter trained on. Right refund decisions went from 70% to 95%. Each run passed 80% compliant, as a ten-step average, by step 57.

step 0step 160 8%93%
decisions flat first 60 steps

Refund calls came later

Wording failures fell from the first step. Refund decisions sat near 70% for about 30 steps, then climbed.

75%67% 6%62% basetrained

Renewals got worse

Plan renewals stayed out of training. Right decisions on them fell from 75% to 67%, since nothing taught the adapter their 14-day window.

training reward0.98 right decisions70% decision check switched off

Unchecked rules stay unlearned

With the decision check off, the training reward reached 0.98. Refund decisions stayed at the base model's 70%.

with the penalty, −0.17 without it, −0.48

Drift needs a brake

Without the base-model penalty, the base model's score for the replies fell from −0.18 to −0.48 per word over 160 steps.