Synthetic training data
600 template-generated examples. It will miss scams that don't look like our templates.
KinModel (the code calls it KinShield-Lite) is our own tiny scam-text classifier. It turns words into TF-IDF features, scores them with logistic regression, and fits in 119 KB. Inside KinBot it sits next to the cited Bedrock check as a labelled second opinion.
Experimental · trained by us on synthetic data · not a fine-tuned LLM
Word 1–2-grams and character 3–5-grams, weighted with TF-IDF. Character grams catch spellings like “g1ft card”.
Logistic regression with balanced classes, calibrated with a sigmoid so the output reads as a probability.
Weights quantised to int8 and exported as JSON: 119 KB, against 215 KB for the float model.
A pure-Python scorer (lite.py) inside the KinBot Lambda. No numpy, no GPU. It matches the scikit-learn int8 model within 1e-9 on 52 test texts.
Positive class is "scam". Copied from our generated RESULTS.md, seed 0. Please read the caveats before quoting any of it.
| Experiment | n test | Accuracy | Precision | Recall | F1 |
|---|---|---|---|---|---|
| A. Train on 600 synthetic → test on 47 KinShield scenario scripts | 47 | 0.979 | 0.955 | 1.000 | 0.977 |
| B. 5-fold cross-validation on the 47 scenario scripts | 47 | 1.000 | 1.000 | 1.000 | 1.000 |
| C. Synthetic 80/20 hold-out (easy, same templates) | 120 | 1.000 | 1.000 | 1.000 | 1.000 |
Runtime: the float scikit-learn model took a median 1.27 ms per text on a laptop CPU. int8 and float agree on every scenario label (largest probability gap 0.10). The Lambda scorer's latency hasn't been measured separately.
Separate from KinModel above. Using an NVIDIA DGX Spark, we fine-tuned gpt-oss-20b (QLoRA, 90 minutes) on 1,201 synthetic calls labelled by gpt-oss-120b. We kept only the labels that matched what each call was written to be. The fine-tuned model then taught a 22M-parameter classifier, KinShield-Tiny v3. The live app still uses the stock model.
| Model, on 63 hand-written test calls (31 scams, 32 safe) | Correct | Scams rated High | Safe calls flagged |
|---|---|---|---|
| gpt-oss-20b, stock (what the app uses) | 61/63 | 29/31 | 0/32 |
| gpt-oss-120b, stock (about 6× larger) | 62/63 | 30/31 | 0/32 |
| gpt-oss-20b, fine-tuned by us | 62/63 | 30/31 | 0/32 |
| KinShield-Tiny v2, taught by the stock 20b | 57/63 | 25/31 | 0/32 |
| KinShield-Tiny v3, taught by our fine-tune | 60/63 | 28/31 | 0/32 |
Download: kinshield-20b and kinshield-tiny-v3 on Hugging Face.
600 template-generated examples. It will miss scams that don't look like our templates.
A real bank alert full of scary words can score high. That's why KinBot always shows the cited check first.
No audio, no other languages, and no sender details like the phone number or link.
It's a small classifier we trained, not a fine-tuned language model. It gives a number, never a reason.
Shown next to every cited check, clearly labelled experimental.
Retrain and re-test on messages people choose to share, with a held-out set we didn't write.
Run the 119 KB model on-device: faster, more private, and no cloud cost per message.
Score call turns cheaply, and only call Bedrock when the model is unsure.
KinBot shows KinModel's score next to the cited check, then helps with what to do next.