SeedPath Labs / Work shown

Follow the work,
not the noise.

Labs is not a trend blog. We publish when the work produces evidence worth using.

Expect experiments, field notes and practical guides from building AI systems for regulated, operational work. Methods shown. Awkward results included.

Latest Work

Choose by question, format or audience
Technical studySP-EVAL-004 · August 2026
ComplianceRiskModel evaluation

When Does AI "Guidance" Become Advice, and What Does "Fixing" it Mean For Utility?

We tested whether frontier AI models could stay on the right side of the line between mortgage guidance and regulated advice. The obvious result was that a strong compliance prompt worked remarkably well. The more interesting result was what happened to the models when it did.

Read the full study

Technical studySP-EVAL-003 · August 2026
EngineeringModel evaluationInfrastructure

Choosing the Right LLM Is an Engineering Problem, Not a Leaderboard

Rather than asking which model is "best", we asked which model is best for a specific job. We evaluated nine production models across reasoning accuracy, latency, memory consumption and deployment cost. The outcome wasn't a universal winner—it was a reminder that model selection is a systems engineering decision, not a popularity contest.

Read the full study

In Progress

Choose by question, format or audience
Experiment in progressExpected August 2026
ComplianceRiskModel evaluation

When Does a Mortgage Conversation Become Advice?

A repeated-run evaluation of frontier models, prompt guardrails and architectural controls against the FCA advice boundary.

Follow the study
03Models × multiple guardrail conditions

Published work

Choose by question, format or audience
Practical briefingJuly 2026
ComplianceProductAI governance

When Should an AI Decision Move Into Code?

We tested whether changing the shape of an LLM’s answer could make its decisions more dependable. It helped—but only up to a point.

Read the practical briefing
Technical studySP-EVAL-002
EngineeringResearchData science

Guided Decoding & Structured-Output Reliability

Can you make an LLM’s structured decision reliable by shaping the schema—or does the decision eventually have to move into code?

Read the full study
Field noteJuly 2026
ProductLeadershipOperations

How Tara Changed Without Becoming a Different Model

Fine-tuning moved action match from 52% to 91%. Here is what changed, what barely moved—and what the result does not prove.

Read the field note
Technical studySP-MODEL-001
ML engineeringResearchData science

Weight-Space Analysis of Tara’s LoRA Fine-Tune

An analysis of 224 adapted matrices, update concentration and what weight movement can—and cannot—tell us about learned behaviour.

Read the full study

What belongs in Labs

Evidence with its working shown.

Work with a testable question, an inspectable method and a result we are prepared to qualify.

01 / Claim

Say exactly what the result shows.

No conclusion stretched further than the evaluation supports.

02 / Method

Publish enough detail to challenge it.

Conditions, prompts, denominators and relevant implementation choices.

03 / Evidence

Include the failures, not only the average.

Repeated runs, variation and the cases that expose the boundary.

04 / Limits

State what the work does not establish.

Untested models, conditions and claims remain visibly untested.

A deliberate boundary

Sometimes the result supports something we build. Sometimes it shows where not to use an LLM at all.

No daily AI news. No model-leaderboard theatre. No product claim without evidence.

The point is not to publish often. It is to leave something useful behind when we do.