When Does a Mortgage Conversation Become Advice?
A repeated-run evaluation of frontier models, prompt guardrails and architectural controls against the FCA advice boundary.
SeedPath Labs / Work shown
Labs is not a trend blog. We publish when the work produces evidence worth using.
Expect experiments, field notes and practical guides from building AI systems for regulated, operational work. Methods shown. Awkward results included.
A repeated-run evaluation of frontier models, prompt guardrails and architectural controls against the FCA advice boundary.
We tested whether changing the shape of an LLM’s answer could make its decisions more dependable. It helped—but only up to a point.
Can you make an LLM’s structured decision reliable by shaping the schema—or does the decision eventually have to move into code?
Fine-tuning moved action match from 52% to 91%. Here is what changed, what barely moved—and what the result does not prove.
An analysis of 224 adapted matrices, update concentration and what weight movement can—and cannot—tell us about learned behaviour.
What belongs in Labs
Work with a testable question, an inspectable method and a result we are prepared to qualify.
No conclusion stretched further than the evaluation supports.
Conditions, prompts, denominators and relevant implementation choices.
Repeated runs, variation and the cases that expose the boundary.
Untested models, conditions and claims remain visibly untested.
A deliberate boundary
Sometimes the result supports something we build. Sometimes it shows where not to use an LLM at all.
No daily AI news. No model-leaderboard theatre. No product claim without evidence.
The point is not to publish often. It is to leave something useful behind when we do.