Not a summary of what someone has said — search already does that. A soul answers the question they never got asked, the way they would have answered it, and tells you how much to trust the answer.
Everything below exists to hit that bar, or to measure whether it was hit. The second half is the part almost nobody builds.
26
subjects
42
soul files built
4.53M
words on disk
0
with a fidelity score
The number that matters, and it has not moved.
The honest state of the project
This page used to say the scoring machinery was finished before any persona was put through it, and that nothing had been measured. Building is no longer the problem. 42 soul files now exist across two extractors, covering 26 people. The number carrying a fidelity score is still zero.
It got worse than that. One soul was scored — 58.3 over four blind probes — and a rebuild by the fork overwrote the score field with nothing. And the one better-powered result the project has produced is a Turing eval the persona failed at 0.95 over 40 samples, where the passing bar is a judge doing no better than chance.
So every profile in the library remains an opinion with citations attached — now in bulk. A page that reported it any other way would be doing exactly what this project exists to catch.
A soul counts as converted only when a schema-valid file exists — and that test has stopped discriminating, because everything passes it now. The distinction that still bites is the next one down: a file that has been measured against held-out material, and a file that has not. Six subjects are documented in detail below because they carry researched corpus notes. The other twenty are counted, not described.
Unrehearsed hours predict fidelity better than corpus size. Marcus shows 0 against 63 items, which is why converting it would produce a reciter whatever it scored.
The pipeline02
One CLI, and the failure each stage prevents
None of the guards are defensive tidiness. Every one is a mistake that was actually made on a real run, usually expensively.
01soul new <slug>
Creates souls/<slug>/ with the corpus, analysis and eval directories.
the failure this prevents
Nothing yet — but the step before it is the one that matters: name the decision you will bring this person. A soul is task-shaped, and “everything X thinks” has no testable output.
The artifact03
The persona is data, not prose
The readable profile is generated from the file, never the reverse. That direction is load-bearing: it is what makes “he believes X” falsifiable in one click.
How they get there. Without this you have a quote machine.
perception
What they notice first, before any belief fires. Get it wrong and every rule fires on the wrong input.
values
Ranked, no ties. Unranked values predict nothing, and every interesting question is a tradeoff.
causal_model
if X → then Y, because Z. 20–40 rules. The layer that answers what the corpus never covered.
evidence_standard
What counts as proof to them. Explains why counter-arguments bounce off.
diagnostic_questions
What they ask before answering. A monologue never has to ask anything.
Every entry carries { video, t, quote }. Below roughly 70% of entries citing the tape, the file is an essay with a schema on top — soul check reports the number rather than trusting it.
How it got here04
Three generations, each built from the last one’s failure
The fourth entry is not a generation. Hermes branched off v4 instead of succeeding it, and both are still in use — it is listed here so the sequence is complete, and taken apart in the next section.
v1
creator-corpus — the essay
The spec was two generations ahead of the implementation. It described six persona documents and named the inference ruleset the highest-leverage one. Neither delivered profile contained one. What shipped was a prose essay written from a word-frequency report — an impression of a person, where the goal is a specification of one.
Channel sweep — Enumerate every tab, transcribe, measure, report coverage.
Source ladder — Rank sources by what they are evidence of, not by how much of them there is.
what it still could not do
Threw away every caption timestamp — one line that cost provenance, retrieval, cadence, drift and speaker attribution at once.
Merged host and guest into one blob, so the tier it called highest-fidelity was the tier it corrupted most.
Required two top-15 unigrams per sentence to detect a mechanism, so a premise stated once — the kind most likely to be load-bearing — was invisible by construction.
v2
The persona becomes data
An essay cannot be checked, scored, routed or diffed. The persona had to stop being prose and start being a file with a schema, generated prose downstream of it rather than upstream.
Timestamps kept end to end — Rolling captions overlap, so a trustworthy pause is a hole between one cue ending and the next starting. That also makes pause-based sentence segmentation possible, which matters because auto-captions are unreliably punctuated — a hole in the clock beats a period.
soul.json — Every entry carries {video, t, quote}. Evidence coverage is measured; below ~70% the file is opinion wearing a schema.
perception, evidence_standard, refusals, eras — Four components v1 had no model of. Perception is the most under-modelled layer in persona work — it decides what the persona attends to at all.
Leak-free eval — Deterministic hold-out by id hash; probes from public metadata, gold from the body the persona never saw.
what it still could not do
No diarization. No council runner. BM25 retrieval only.
v3
The living-person layers
v2 modelled a person as a reasoner — perception, values, rules, voice. A reasoner is a stateless function: same question in, same answer out, forever. Three things were missing, and none of them were more beliefs.
Modes — State. The persona picks the register the question would put them in, not the corpus average.
Relationship memory — Continuity. Open commitments get asked about first — “you said you'd start three weeks ago, did you?” Advice escalates rather than repeats, and corrections never recur. The highest felt-aliveness per line of code in the system.
Biography and the enemy — Origin. Every belief system is a reaction to something specific. Name what they argue against and you can predict positions the corpus never touches.
The wound — The unresolved version of the enemy. When a question lands there the real person gets hotter and less rational, and a persona that stays measured across its own wound is the most common tell.
Turing eval — Strictly better than lexical F1, which rewards vocabulary rather than being indistinguishable.
what it still could not do
No soul has been converted yet. Six exist as prose.
Diarization still manual — interviews remain the highest-value and most corrupted tier.
Composite advisor and council runners specified, not built.
Voice and embodiment deferred by choice.
hermes
The fork that builds, and stops measuring
v4 could specify a person and score one, and in practice did neither at volume — one soul was ever put through the eval. The fork inverted the tradeoff on purpose: strip what makes a run slow or cautious, build the whole library, and accept a personal-use tool rather than a defensible one. It is a strict superset of v4's scripts, plus six more. Twenty-five souls came out of it against v4's seventeen, and not one carries a score.
Data-driven build_soul.py — v4 filled the soul file from keyword-matching templates, which is how a domain came out as the string “girl, girls, good, life”. The fork reads analysis.json — biography, enemy, mechanisms, directives, contradictions, drift, negative space — so the file describes the corpus rather than the vocabulary.
Evidence-backed voice — v4's voice section was 0.2 KB and cited nothing, which is the section most able to make a wrong soul persuasive. The fork searches each lexicon term against the mechanisms and directives for a real citation: 8 of 8 evidenced.
affect, diagnostic_questions, drift, the_enemy — Four sections v4 either lacked or left as stubs. The soul file roughly doubles — 117.8 KB against 49.1 KB — and the added weight is structure, not prose.
render_soul and render_transcripts — SOUL.md and one readable transcript per video, timestamps clickable. The human artifact is generated from the file, which keeps the direction v2 established rather than quietly reversing it.
organize search — Ranks which souls in the library actually cover a topic. The first thing in the project that treats the library as a library rather than a folder.
--slim — Deletes raw .info.json and .vtt after extraction, and consolidates the subs/src duplicate that had been writing every caption file twice. Freed 4 GB; a typical soul went 35 MB to 4 MB.
what it still could not do
It drops the fidelity block when it rebuilds a soul. The one measured soul in the project lost its score this way.
Zero of its 25 souls have been scored, so the superset claim holds for scripts and not for evidence.
It is a jailbroken personal-use fork — the emulation restrictions v3 specified are reduced to a single rule about absent third parties.
Three of its 25 souls have no rendered SOUL.md, so the readable artifact silently lags the machine one.
The fork05
Hermes — a superset in scripts, a regression in evidence
tools/soul-extractor-hermes branched off v4 rather than succeeding it, and both are still run. It has every script v4 has plus six more, writes to souls/hermes/<slug>/, and is driven by bin/hermes-soul. It produced twenty-five souls against v4’s seventeen in a fraction of the time.
It is also a deliberately jailbroken personal-use tool. The emulation restrictions v3 specified are reduced to one rule: a soul will not give first-person advice about a specific named real person who is absent from the conversation. That is a smaller guarantee than v3 was designed to make, and it is the trade the fork exists to take.
axis
v4
hermes
Soul file content
v4Keyword-matched templates. 49.1 KB.
hermesRead from analysis.json. 117.8 KB.
Voice
v40.2 KB, no citations.
hermes5.5 KB, 8 of 8 terms evidenced.
Human artifact
v4Hand-written prose profile.
hermesSOUL.md generated from the file, timestamps clickable.
hermesNone scored, and a rebuild drops the score field from any soul that had one.
The last row is the one that matters. A fork that builds faster than it measures makes the library bigger and the project weaker, and this one deletes the score field of any soul it rebuilds — which is how the only real measurement in the project was lost without an error.
From the 5,486-item run06
What only shows up at scale
A fifty-item corpus surfaces none of these. Most tooling is never run hard enough to find out how it fails, so it accumulates behaviours like these and nobody notices — the output always looks like an answer.
A rebuild silently deleted the only fidelity score in the project
silent failure
marcus-hullaster was scored on 2026-08-27 — 58.3 across four blind probes, written into the soul file's fidelity block. The hermes fork then rebuilt the soul from the corpus, and its writer emits a fidelity block containing only evidence coverage. The score, the probe count and the evaluation date were overwritten with nothing. No error, no warning, and the file still validates, because a fidelity block with one field is a legal fidelity block. The eval artifacts under eval/ were untouched, so the measurement survives on disk while the file that is supposed to carry it does not — which is the worst arrangement of the two, because the soul file is the thing that travels.
The Turing eval was run, and the persona failed it decisively
silent failure
The design's strong pass is a judge doing no better than chance at telling the real person from the persona — around 0.5. Over 40 paired samples the judge scored 0.95, catching the generated side 95% of the time. That is not a near miss; it is the persona being obvious. It also reframes the fidelity number: 58.3 came from 4 probes, and this came from 40, so the better-powered result is the one saying the thing does not hold up. Every soul in the library was built by the same pipeline that produced it.
n=40, accuracy 0.95, generated_caught 0.95.
Being blocked looked exactly like having no captions
silent failure
The fetcher shelled out to yt-dlp without --ignore-no-formats-error, so format resolution ran even under --skip-download, and format resolution is what trips YouTube's bot challenge. The 456 failures were recorded as no_captions — which reads in the manifest as “this creator has no captions” rather than “we got blocked”. For a tool whose entire job is corpus completeness, that is the worst available failure mode: it does not error, it lies quietly and consistently.
59 of 515 transcribed (11%). After adding two flags, 508 of the same 515 (98.6%).
A captured batch with no timeout is a blind hang
cost real time
Every URL went to a single yt-dlp process with output captured, no timeout and no per-item progress. It sat for ten minutes producing zero files and zero signal. The instruction to report corpus size before going deep is unfollowable when nothing is observable. Fixed with one subprocess per item, a 120s timeout, a thread pool, and a counter every ten items.
Six workers, not ten
method
Caption fetching is latency-bound, not CPU-bound. Serial would have taken 43 minutes for 515 items. Six workers took 475 seconds. Ten workers get throttled.
43 min → 7.9 min.
The scope conversation is about reading, not downloading
method
Enumerating first cost seconds and answered the question up front: 515 videos, 217.5 hours. The old rule said agree a scope above ~50 hours — but the right call was to take all of it, because fetching is cheap and the scarce resource is analysis context, not disk.
Measure before theorising — and be willing to be wrong
method
The cross-platform divergence section was drafted off a nine-file Shorts sample and came out backwards. With all 4,886 Shorts in, the real split inverted: Shorts are the motivator, long-form the operator, interviews the reflector. The rule that says read the frequency tables before forming a thesis was correct and got violated anyway, so the guardrail is now stronger — do not write the divergence section until every tier is 100% complete.
The mid-catalog slice was the highest-yield read
method
4.25M words does not fit in one context. Five parallel reads over disjoint, purposefully chosen slices — top-20 by views, a random mid-catalog sample, two interview halves, a voice-mechanics pass — each returned a structured report. The random mid-catalog slice beat all of them, and a single-context pass would never have read it: it is where the ideas that never made a hit video live.
One interview in twenty-two was a re-cut of two others
cost real time
About 70% verbatim overlap. Cheap to catch with a shingle test, expensive to miss, because a duplicate inflates every frequency count in the tier it lands in.
Standing rules07
What this is allowed to be
01A soul is a model of a person, and the person is not consulted.
02Reconstruct the generator, not the transcript. If it can only answer what its subject already answered, you built a search index.
03Voice last. Sounding right is what makes a wrong soul persuasive.
04Every load-bearing claim is [quoted] with a timestamp, or [derived] with the rule ids. Derived reasoning is never presented as something they said.
05Souls advise. They do not decide, and they do not vote.
06Corpora never leave the machine. The soul file travels; the tape does not.
07A soul without a fidelity score is an opinion. Forty-two of forty-two currently qualify.
Sources include material held lawfully because I was in the room — mentorships, programs, calls, notes. That material stays local, stays out of version control, and never ships. Someone else’s paywall is out of scope; being in the room is what makes the material yours to study.
A soul is a model of a person, and the person is not consulted✦Reconstruct the generator, not the transcript. If it can only answer what its subject already answered, you built a search index✦Voice last. Sounding right is what makes a wrong soul persuasive✦Every load-bearing claim is [quoted] with a timestamp, or [derived] with the rule ids. Derived reasoning is never presented as something they said✦A soul is a model of a person, and the person is not consulted✦Reconstruct the generator, not the transcript. If it can only answer what its subject already answered, you built a search index✦Voice last. Sounding right is what makes a wrong soul persuasive✦Every load-bearing claim is [quoted] with a timestamp, or [derived] with the rule ids. Derived reasoning is never presented as something they said✦