Lineage

Docs · view on GitHub

How to use Lineage

Lineage is a cooperative family documentation and research platform. A family collects its material: recordings, photographs, documents, letters, scans, links. Lineage turns it into a researched, sourced, browsable record: the people, places, events, vessels, organizations and objects in it, a genealogy built from evidence, a timeline, and an encyclopedia of all of it (the Familypedia), with every fact linked to the source that supports it. Stories, a printed book and narration are exports of that record, made at the end.

This guide follows the work in the order a family usually does it:

Part What happens
1. Set up a lineage install the tools, start a project, open the dashboard
2. Invite contributors the people who add material, and the review of what they add
3. Add sources intake of every file; transcription of recordings; photographs catalogued
4. Research records the family's own papers and public records, catalogued by type and snapshotted
5. Build the record the timeline, the genealogy and the Familypedia
6. Make stories and a book story units, the chapter map, writing, building, printing, narration
7. Troubleshooting · 8. Honest limits

It uses the sample project throughout: an invented family, the Calders. Ruth Calder (born 1938 in Duluth, Minnesota) is interviewed by her grandson Sam, and she talks about her grandfather Anders Calder, who came from Norway. Everything about them is made up, so you can open every file and see exactly what each stage produces.

Four things hold everywhere in Lineage, and the rest of this guide is how they are kept: - Nothing is invented. A fact enters the record only from a source, and it carries a citation back to that source: a timestamp in a recording, a record's catalogue number, a URL. - Evidence has tiers. Every fact is witnessed (the speaker saw it), told (someone named told them), lore (the family's story, source unclear) or documented (a record shows it). The tier travels with the fact into every view and every export. - Contradictions are kept. When two tellings disagree, or the family and a record disagree, both stay, side by side, and the disagreement is shown. Nothing picks a winner silently. - A person decides. Machines propose: tags, genealogy links, chapter maps, connecting sentences. People accept them. The places where a person must decide are named, and the tools stop there.


1. Set up a lineage

1.1 What you need

The material. Whatever the family has. Recordings in any common format (m4a, mp3, wav, aac, flac, aiff, ogg, opus, wma, mp4 or mov; phone voice memos are fine), photographs and scans, PDFs, letters, Word documents, notes, links. Originals are only ever read, never changed.

A Claude subscription and Claude Code. The model work (understanding sources, the timeline, the genealogy rebuild, story units, shaping, captions, the introduction) is done by Claude Code following the rules in .claude/skills/. You talk to it in plain English ("where are we?", "build the timeline", "shape the units for chapter 3") and it runs the scripts and writes the files. Text is sent to Anthropic while it works; audio is not (see PRIVACY.md).

A Mac or Linux computer. Windows is not tested. Apple Silicon works, but speech recognition runs on the CPU there, because the engine has no Metal backend. An NVIDIA GPU on Linux makes it much faster.

Software:

Tool Why Install
Python 3.10+ everything, including the dashboard usually present; on Debian/Ubuntu also apt install python3-venv
ffmpeg reads recordings brew install ffmpeg / apt install ffmpeg
poppler (pdfinfo, pdffonts, pdftotext) text from PDFs at intake; page counts and the print preflight brew install poppler / apt install poppler-utils
tesseract (optional) OCR of scans and image-only PDFs at intake brew install tesseract / apt install tesseract-ocr
WhisperX, pyannote, PyTorch speech recognition and speaker labels for recordings make install-transcribe (pinned versions)
Typst rendering stories as pages, and the printed book brew install typst, or see Typst's install page

Fonts. EB Garamond is bundled in fonts/ under the SIL Open Font License, and every build points Typst at it. You don't need to install anything. The one exception is the map tool (section 6.6), which uses EB Garamond only if it's installed system-wide.

Disk and network. The first transcription downloads about 5 GB of models. Each hour of audio needs about 110 MB more for a working copy. The first book build downloads one small Typst package (droplet, for the drop caps). After that, everything except the model work and the research runs offline.

A Hugging Face token (only if you have recordings). Hugging Face is where the speech models are published. The speaker-labelling model (pyannote) is gated: it's free, but its authors ask you to accept their licence and share contact details before you download it. The token is how the download proves you've done that. The model runs on your computer. The token is only used to download it, and no audio is ever sent. To set it up once:

  1. Create a free account and a Read token at https://huggingface.co/settings/tokens.
  2. While logged in as that same account, open both of these pages and accept the conditions on each: - https://huggingface.co/pyannote/speaker-diarization-3.1 - https://huggingface.co/pyannote/segmentation-3.0

The first model depends on the second, and accepting only the first is the usual mistake. 3. Put it in your shell, never in a file in the project:

export HF_TOKEN=hf_xxxxxxxxxxxxxxxx

Without a token you can still transcribe. The speakers just stay unlabelled until you run the diarization step later.

1.2 Install the tools

git clone https://github.com/rexsaurus/Lineage.git ~/Lineage
cd ~/Lineage
make install              # Python venv in .venv with the build tools; checks for typst
make install-transcribe   # adds WhisperX + pyannote to the same venv (a large download)
make sample               # optional: proves the build works (see the README)

1.3 Start a lineage

A family's lineage lives in its own folder, outside the repo, so your family's material never mixes with the public code:

make new PROJECT="$HOME/lineages/calder"

Use $HOME rather than ~. In zsh, the default shell on macOS, PROJECT=~/lineages/calder is not expanded, and make would create a folder literally named ~ inside the repo.

make new creates:

calder/
  book.yaml            names, the subject's birth year, print settings (fill this in now)
  lineage.lock         the Lineage release this project runs (section 1.5)
  Makefile             make dashboard, make lineage-version, make update-lineage
  CLAUDE.md            standing rules Claude Code reads at the start of every session
  local-overrides/     the few things specific to this family (section 1.5)
  .gitignore           keeps audio, photos, raw web caches and PDFs out of git
  .claude/skills  ->   a link to the skills in your Lineage checkout
  audio/               recordings (never modified)
  transcript/          raw JSON, verbatim and clean transcripts, sessions.csv, corrections.json
  facts/               timeline.csv, glossary.md, gaps.md (questions for you), records/
  content/units/       one file per story
  data/                archives.csv, chapters.csv, reader_glossary.csv, indexes
  chapters/            generated chapter files
  book/front/          title, copyright, dedication, contents, introduction
  photos/source/       original images (never modified)
  photos/print/        print-ready copies
  output/              PDFs and spreadsheets

The dashboard adds its own files as you use it: sources/ (files added through the Sources tab), lineage.json (settings), data/sources.json (the source index), data/genealogy/, data/familypedia/, data/requests.json, and .lineage/ (logs, rendered pages, thumbnails). app/README.md lists everything it reads and writes.

1.4 Fill in book.yaml

book.yaml is the project's identity file, and every skill reads it first. (It keeps its name from when the book was the only output; the dashboard's Family details page holds the family-level identity: title, family name, summary, crest.) Here is the sample's:

title: "The Lake Was Always There"
subtitle: "A Life of Ruth Calder"
narrator:                       # the main speaker on the recordings
  name: "Ruth Calder"
  label: "Grandma"              # speaker label used in the transcripts
  birth_year: 1938              # anchors every "when I was twelve"
  birthplace: "Duluth, Minnesota"
interviewer:                    # the family member who recorded the interviews
  name: "Sam Calder"
  label: "Sam"
other_speakers: []
family_figures:                 # relatives who get their own family-history section
  - "Anders Calder"
proper_nouns:                   # KEEP SHORT: surnames and odd place names only
  - "Calder"
  - "Duluth"
transcription_locked: true      # set once timestamps are cited (section 3.2)
print:
  trim: "7x10"                  # 6x9 | 7x10 | 8x10 | 8.5x11
  color: "bw"                   # bw | color
  printer: "kdp"                # kdp | ingramspark | lulu | blurb
  in_chapter_contents: false
narration:
  mode: "third_person"
  subject_name: "Ruth"          # how stories refer to the subject after first mention
front:
  introduction_title: "Introduction"

The field that matters most is birth_year. People date their lives by age: "when I was twelve", "the year I started school", "after I turned sixteen". Every one of those becomes a year by adding it to the birth year. Get it from a document if you can. If it's off by one, every derived date in the timeline, the Familypedia and the book is off by one.

Two more to get right early: - label values are the names that will appear on transcript paragraphs (**Grandma**, **Sam**). They must match what you put in the speaker map in section 3.2. - proper_nouns are a few names the speech engine will otherwise mangle. Keep the list short (section 3.2 explains why).

1.5 lineage.lock and updates

A project uses Lineage; it doesn't contain it. It keeps only its own material (sources, transcripts, stories, people, records, photos, settings). Everything else (skills, style rules, templates, scripts, the dashboard) stays in Lineage and reaches the project as a release. lineage.lock pins the release the project runs:

make lineage-version          # what this project runs
make update-lineage           # what a newer release would change, and which of your files it touches
make update-lineage APPLY=1   # install and pin it
make update-lineage TO=v0.4.0 APPLY=1   # a particular release instead of the newest

Releases install read-only under ~/.lineage/releases/<tag>. An update never touches your material. When a release changes something that would alter stories already generated (the book template, fonts or skills), those stories are marked stale and listed, never rewritten.

Improvements found while working on one family's record go into Lineage and come back down as a release, so every project gets them. Things only one family would want go in that project's local-overrides/.

1.6 The dashboard

make dashboard                     # in a project: the dashboard of the pinned release
~/Lineage/app/lineage ~/lineages/calder   # or from a checkout, on any project folder

It runs on your own computer at http://127.0.0.1:8777 and binds to 127.0.0.1 only. python3 server.py --demo (in app/) shows invented sample content with a banner; it is never the default.

There is no wizard. Every tab works whenever you open it and says plainly what it still needs:

Tab What it's for
Home story of the day, a featured relative, Needs you (one-click actions ordered by what they unblock), Request more (question lists built from open questions, gaps and unconfirmed links), counts, and an activity feed
Sources every file the family has added, and its intake (section 3.1)
Familypedia an article for every subject the material names (section 5.3)
Genealogy the tree, derived from the sources with evidence on every link (section 5.2)
Stories the written stories, rendered as pages, with narration (sections 6 and 6.10)
Timeline every dated event, with tiers, conflicts and gaps (section 5.1)

Behind the settings gear: Family details (title, family name, summary, crest), Connectors (Google Drive, GitHub, Anthropic, OpenAI, ElevenLabs, agent CLIs), Contributors (section 2) and Project settings (repo, Drive folder, narration voice, writing style, story templates, trim and printer, the terminal command). A dot on the gear means something needs attention.

The Terminal button in the header (or Ctrl+`) opens a drawer over any tab with a real terminal running Claude Code in the project: the Genealogist, quick prompts, the pipeline actions and the approvals list. Buttons that need real work done ("Generate", pipeline actions) hand it to the Genealogist there.

1.7 Every session: set up the shell, then open Claude Code

export LINEAGE="$HOME/Lineage"     # put this line in your shell profile
source "$LINEAGE/.venv/bin/activate"     # so `python` is the Lineage venv
cd "$HOME/lineages/calder"
claude

Claude Code picks up the skills from .claude/skills/ and the rules from CLAUDE.md. Start each session by asking "where are we?". It runs the status check and resumes at the first unfinished stage. You can run the same check yourself:

make -C "$LINEAGE" status PROJECT="$PWD"

The output lists the pipeline's stages, each marked ✓ or ·, followed by a NEXT: line.

Everything in this guide can be done more than one way: in the dashboard, by asking Claude ("transcribe the recordings", "build the timeline", "rebuild the genealogy"), or by running the commands shown here yourself. The commands assume you are in the project folder with the venv active.


2. Invite contributors

A lineage is cooperative: many people add to one shared record. The family is the subject; contributors are the people who add material about it, from wherever they are.

Settings → Contributors holds: - people and roles: contributor (adds material), reader, editor; - invite links, one per person, with an expiry, revocable; - requests outstanding: what you've asked each person for (Home's Request more builds these lists from open questions, gaps and unconfirmed links, and saves them as asked or answered in data/requests.json); - the shared folder contributors can drop files into (with the Google Drive connector); - the review queue: material a contributor added waits here until someone accepts it. Home's Needs you says when something is waiting.

Members and invites are kept in .lineage/family.json in the project.

What is not there yet. The dashboard runs on your own computer, so an invite link only works on that machine until the project is hosted, and the page says so. The management side (people, roles, invites, requests, the review queue) is built; the contributor-facing view behind an invite link waits for hosting (see app/ROADMAP.md). Until then, the practical route is the shared Drive folder, email, or the post: you add what arrives through the Sources tab and record who it came from.


3. Add sources

3.1 The Sources tab: intake

Drop files onto the Sources tab, or put them in the project and they are indexed where they are. Every file runs the same visible stages, and each stage's result is kept in data/sources.json:

Stage What happens
Saved (ingest) the file is copied into sources/ and given a SHA-256 hash. A file already in the project is recognised by its hash and not added twice. The original is never modified or renamed.
Reading (extract) text comes out: a PDF's text layer, OCR for scans and image-only PDFs (needs tesseract), Word and text files directly. With an Anthropic key, an image also gets a short description and a transcription of any writing on it. It never names a person from how they look; a person is named only when writing on the item names them.
Understanding a summary, the people, places and organizations named, a date range and the kind of item (letter, photo, certificate, record, transcript, recording, notes). With an Anthropic key this is a model pass; without one, a simpler heuristic. Everything here is marked as derived.
Indexed the source joins the full-text search (data/search_index.json).
Drive (optional) a copy goes to the project's Drive folder when the Google Drive connector is enabled and a folder is set.

A stage that fails never loses the file; it shows the failure and can be re-run. Your edits in the edit drawer (name, summary, people, dates, notes) are kept apart from the derived values and always win, including across a re-ingest. Sources can be renamed, trashed and restored (and then deleted permanently), and most actions work in bulk.

Tagging a source to the people and subjects it concerns happens in the Familypedia (section 5.3). A gallery and lightbox for photographs, and annotation and people tagging on the image itself, are not built yet (app/ROADMAP.md, #5).

3.2 Recordings: transcription

A recording becomes useful when it becomes a timestamped, speaker-labelled transcript: then every fact taken from it can point at the second it was said. A recording added through the Sources tab shows waiting for transcription until its session's transcript exists; re-ingest it afterwards and the transcript text becomes its searchable content.

Transcription runs on files in audio/. Copy the recordings there by any means. Then:

$LINEAGE/scripts/transcribe.sh --dry-run   # list sessions and the plan; do nothing
$LINEAGE/scripts/transcribe.sh --smoke     # the shortest session only: check it all works
$LINEAGE/scripts/transcribe.sh             # everything, shortest first

For each recording, this: 1. gives it a session ID (S1, S2, … in recording-date order) in transcript/sessions.csv. IDs never change once assigned, so adding a recording later never renumbers anything. If a file has no embedded date, fill in the recorded column by hand; it won't be overwritten. 2. makes a 16 kHz mono working copy in transcript/work/; 3. runs WhisperX (large-v3, with word timestamps) into transcript/raw/S1.json; 4. if HF_TOKEN is set, adds speaker labels (diarization).

It is safe to stop and restart. A session that already has its JSON is skipped, and output is only moved into place when a session finishes, so a crash never leaves a half-written file that looks done. A log is kept in transcript/work/transcribe.log.

Other useful forms: transcribe.sh S3 (one session), transcribe.sh --no-diarize (speech only, label speakers later). Environment overrides include MODEL, LANGUAGE (default en), DEVICE, and MIN_SPEAKERS/MAX_SPEAKERS. By default it expects two speakers plus any other_speakers.

Keep the initial prompt short

The speech engine accepts an "initial prompt" of vocabulary to listen for, and transcribe.sh builds it from proper_nouns. A long prompt gets echoed back into the transcript as fake speech: during a pause, the model "hears" your list of names read aloud in the middle of a real answer. The script warns above 120 characters. Keep it to a few surnames and unusual place names, and fix every other misspelling afterwards (below).

After every run, search the output for your prompt text. If an echo slips through, cut it at render time in transcript/corrections.json: - drop_segments.patterns drops a whole raw segment that is nothing but echo; - drop_word_runs.runs removes an echoed phrase word by word from inside a real segment; - scrub_inline.patterns strips an echo embedded inside a real paragraph.

Long recordings: chunked diarization

Speaker labelling gets disproportionately slower as files get longer. Measured on a CPU, it took about 32 seconds of compute per audio-minute on 6–10 minute files but about 119 seconds per audio-minute on a 34-minute file. A two-hour interview in one piece would take something like 14 hours. So any session longer than 20 minutes is automatically labelled in overlapping 10-minute windows, which are then stitched together. You can also run it directly:

python $LINEAGE/scripts/diarize.py                    # every session without speakers
python $LINEAGE/scripts/diarize.py S1 S3
python $LINEAGE/scripts/diarize_chunked.py S5         # force chunking
python $LINEAGE/scripts/diarize_chunked.py S5 --chunk 600 --overlap 60

Each window logs a line like stitched 2/2 by overlap. Anything less than all speakers matched deserves a listen at that seam.

Never re-run speech recognition once timestamps are cited

Two runs over the same audio never produce the same segment boundaries. Every citation in the record ([S2 00:14:07]) would silently point at the wrong words. So:

GATE 1: confirm the speakers

Diarization labels voices SPEAKER_00, SPEAKER_01, … and the numbers mean nothing. Listen to a minute of each session and decide which cluster is the subject. The interviewer usually asks short questions, and the subject tells long stories. Then write the map into transcript/corrections.json, using the labels from book.yaml:

{
  "speaker_map": { "SPEAKER_00": "Grandma", "SPEAKER_01": "Sam" },
  "spelling": [ { "find": "\\b[Cc]aulder\\b", "replace": "Calder" } ],
  "drop_segments": { "patterns": [] },
  "scrub_inline": { "patterns": [] }
}

Then render:

python $LINEAGE/scripts/render_transcripts.py

This writes three layers: - transcript/verbatim/S1.md: every word, with low-confidence words marked [?word?]; - transcript/clean/S1.md: fillers (um, uh) and stutters removed, everything else kept. This is the layer everything quotes and cites; - transcript/master.md: all clean sessions in order.

A clean paragraph looks like this:

**Grandma** [S1 00:00:29] When I was little, maybe seven or eight, I'd walk my dad's lunch down to the ore dock. A tin pail. ...

The render prints a line per session ending in unmapped none when every voice has a name. If Claude is running this stage, it stops here and shows you three sample turns per speaker to confirm. Nothing is written until you say the mapping is right.

One limitation: render_transcripts.py applies a single speaker_map to every session. Cluster numbers can come out swapped from one session to the next. If that happens, you can render that session by hand with the skill's converter, but this skips the spelling and echo fixes, and a later full re-render will overwrite it:

python .claude/skills/interview-transcriber/scripts/whisperx_to_md.py transcript/raw/S3.json \
  --session S3 --map SPEAKER_01=Grandma --map SPEAKER_00=Sam \
  --verbatim transcript/verbatim/S3.md --clean transcript/clean/S3.md

Also confirm the basics in book.yaml (names, birth year) at this gate. Claude keeps a spelling list in facts/glossary.md. Each spelling you approve goes into corrections.json → spelling, and you re-render.

3.3 Photographs: catalogue with evidence

People will treat a photograph's label as fact for generations. So every label says what it rests on, every guess looks like a guess, and nothing that isn't a photograph can pass for one. The photo-processor skill holds the full rules.

A family photos document (Word, PDF, an exported Google Doc, a zip or a folder of scans) is split into catalogued images:

python .claude/skills/photo-processor/scripts/extract_photos.py "Family photos.docx"   # or a PDF, zip or folder

Images get permanent IDs (P001, …). Originals are copied to photos/source/ and never edited, and any nearby text is saved as original_caption, which is often the best evidence there is. Export a Google Doc as .docx first. Everything about each image goes in photos/photo_index.csv (which the Familypedia reads), and each identification records its basis, strongest first:

  1. inscription: writing on the photo or its back, a printed lab date;
  2. document: a caption in the family's photo document, an archive catalogue, the owner's word;
  3. transcript: the subject describes this scene (cite [S2 00:31:05]);
  4. visual estimate: clothing, cars, print format. Never better than medium confidence.

Never name a person from facial resemblance. A wrong name in the record is worse than none. Without an inscription, document, transcript or the owner's word, describe instead ("unidentified woman, about 30"). Dates are exact only if inscribed; otherwise they're a range or "about 1925". Only you or the owner can mark an image confirmed.

Generated images are marked as generated everywhere: rows in photos/photo_index.csv whose subject starts "Illustration" show as illustrations in the Familypedia, and section 6.6 sets the rules for using one in a book.


4. Research records

The family's version is half the record. The other half is what the archives hold: census returns, military rosters and pension files, ship registers and crew lists, digitized newspapers, museum and library catalogues, period books. Lineage catalogues both, snapshots what it finds, and sets the records beside the family's version. The records-archives skill does this work; ask Claude to "research Anders Calder" or "index the family's papers".

4.1 The family's own papers

data/archives.csv is an index of what the family holds and what it has lost: Bibles, letters, discharge papers, albums, the recordings themselves, heirlooms. Each row is typed (photographs · letters · documents · bible · military · legal · recordings · heirlooms · institutional · other), says who holds it in general terms, and has a status (confirmed · mentioned · unknown · lost · destroyed · institutional). Lost and destroyed items get rows too: "the letters burned in 1962" saves a future relative years. Street addresses and phone numbers live only in the private column, and a living holder is printed only with print_permission: yes.

python .claude/skills/records-archives/scripts/archives_tools.py check   # flags addresses/phones in public columns
python .claude/skills/records-archives/scripts/archives_tools.py asks    # follow-up list, one call per relative

4.2 Public records

Research beyond the family's own material is welcome anywhere in the record. Check the records first, because archives often hold the person.

python $LINEAGE/scripts/fetch_records.py anders-calder https://example.org/roster/page-214
python $LINEAGE/scripts/export_sources.py --dry-run
python $LINEAGE/scripts/export_sources.py --restricted '^museum/'

This copies text files to facts/records/sources/ and writes MANIFEST.csv listing every raw file with its size and SHA-256 hash, plus whether it was exported and why not. Restricted paths are never copied. Rules can live in facts/records/export.yaml; see the script's header.

Low-confidence OCR from a record goes into the fact sheets flagged as such, and never into anything written without being checked.

A dashboard screen for pasting record links and scanning them into the typed catalogue, with a by-archive view, is not built yet (app/ROADMAP.md, #7). Today the research runs through Claude Code and the files above.

4.3 Routes

Journeys (a crossing, a voyage, a regiment's march) are kept as sourced CSVs, one row per recorded point:

seq,date,place,lat,lon,kind,aboard,source,label,gap_before
1,,Bergen,60.39,5.32,port,subject,<record ID or URL>,Bergen,
2,,New York,40.70,-74.01,port,subject,<record ID or URL>,New York,yes
3,,Duluth,46.78,-92.10,port,subject,<record ID or URL>,Duluth,

Only seq, lat and lon are required, but every point should carry a source. A file named *track*.csv or *route*.csv under facts/ gives the Familypedia its ports, positions and map coordinates. A new leg, gap_before, or a row without coordinates is an unrecorded leg, and it is shown as unrecorded rather than drawn. Section 6.6 turns the same file into a printed map.

4.4 When records contradict the family

Keep both, and say so. Never silently correct the family's story, and never silently repeat it. The timeline keeps both tellings (section 5.1), the genealogy rebuild keeps contradictions for review (section 5.2), and the Familypedia shows the family's tiers and "what the records show" as separate sections of the same article. Log the conflict in facts/gaps.md, where it becomes a question for whoever can answer it. When the disagreement reaches a written story, section 6.5 says how to tell both.

Say honest gaps out loud ("Which route the boat took that winter, no surviving record says") and never fill them. Don't give an ancestor a famous battle or ship their unit or crew did not have.

Include a note on the name whenever a record could be confused with a similar one, such as two ships or two men with the same name: say which is which, and why, and keep the reasoning in facts/records/<person>/. The worked example (EXAMPLE-CHAPTER.md) ends with exactly such a note about two whaleships called Hannibal.


5. Build the record: timeline, genealogy, Familypedia

5.1 The timeline

People don't tell their lives in order. facts/timeline.csv puts every event on one line, and each date shows how it was worked out. Ask Claude to "build the timeline" (the timeline-organizer skill), then check its work:

python .claude/skills/timeline-organizer/scripts/timeline_tools.py sort
python .claude/skills/timeline-organizer/scripts/timeline_tools.py check   # fix whatever it reports
python .claude/skills/timeline-organizer/scripts/timeline_tools.py md      # facts/timeline.md, by decade

Resolving relative dates

Each row records the arithmetic in date_basis, and how it will be said in date_display. From the sample:

event quote date_basis date_display confidence
E003 Ruth walks her father's lunch to the ore dock "maybe seven or eight" 'maybe seven or eight' + birth year 1938 about 1945 medium
E004 The big snow; school closed for a week "around 1950, I think" 'around 1950, I think' around 1950 medium
E005 Ruth begins nursing school "nursing school in 1956" stated 1956 high

The rules: - Never invent precision. There's no month or day unless she said it or a document gives it. - Hedges carry over. "Around 1950, I think" stays "around 1950". Nothing downstream may state a date more precisely than date_display. - "Right after the war" is resolved by saying which war and why, labelled as historical context. - A bare "yeah" to a leading question ("Was that 1952?" "Yeah.") gets low confidence.

Conflicts are kept, not settled

When two tellings disagree, or the family's story and a document disagree, both stay. The conflicts column says what disagrees with what, confidence drops to low, and a question for you goes into facts/gaps.md. Nobody picks a winner silently.

The Timeline tab

The dashboard draws the same file as a vertical spine with decade bands and a year rail. Each card shows the date and its precision, the tier, the people, the place, the citations and links to stories. Conflicts get their own cards; gaps get cards with "Add to questions"; events with no date wait in an undated drawer. Filter, search, star, and export as SVG, PNG or a printable appendix. (The printed book also takes an appendix, "A Timeline", built by make draft from the high- and medium-confidence rows.)

5.2 The genealogy

The Genealogy tab derives the family tree from the sources. Every link carries its quoted evidence: no evidence, no link.

5.3 The Familypedia

The Familypedia is the record's encyclopedia: an article for every subject the material names, of nine types:

person · place · event · vessel/vehicle · organization/unit · object · publication · occupation/trade · theme

Each article has a lead from the material, a typed infobox, tier sections (witnessed · told · lore · what the records show), the passages that mention it, its sources with thumbnails, typed records with archive, number, link and retrieval date, photographs and marked illustrations, stories, related articles, backlinks, open questions, and a separate "Beyond the family" section for public background. Browse by type, A–Z, most material or needs more; search the full text with type filters; follow [[links]] across types.

Where articles come from. Only the project's files; everything is optional, and a project with less material simply has fewer articles. The main inputs:

Input Gives
content/units/*.md front matter people, places, subjects ("vessel: Hannibal") subjects, and the passages that mention them (section 6.2)
facts/timeline.csv events, their people and places, tiers, conflicts
facts/people/*.md (# Name, Also called:, Relationship to …:, Dates:) person profiles and other names
knowledge/graph.json (or nodes.csv + edges.csv) typed subjects, records and relations (crew_on, master_of, served_in, held_by, …)
data/archives.csv, facts/records/**/sources.csv the records catalogue and research sources (section 4)
facts/**/*track*.csv, *route*.csv ports, positions and map coordinates (section 4.3)
photos/photo_index.csv photographs, with illustrations marked (section 3.3)
facts/gaps.md, facts/records/**/context_*.md open questions; public background for "Beyond the family"

app/README.md has the full table, including where map coastlines come from. No map tiles are ever fetched: the map is drawn from the records' own coordinates, with unrecorded legs shown as unrecorded.

Tagging. Any source, record, photograph or event can be tagged to any article, one at a time or in bulk from the Records and Photographs views. Lineage suggests tags only from names written in the item, shows the words that name them, and tags nothing until you accept: no faces, no resemblance. Tags, your notes, infobox values, "same as" merges and subjects you create are saved under data/familypedia/, and they are yours: rebuilding from the material never overwrites them.


6. Make stories and a book

Everything above is the record. This part turns it into stories (written pieces, which the dashboard renders as pages and can narrate) and a printed book. It is an export: the record is complete without it, and the book can be rebuilt from the record at any time.

The book pipeline adds two more points where a person decides, after GATE 1 (section 3.2):

Gate When What you decide
GATE 1 after transcription which voice on the tape is the subject and which is you; the basics in book.yaml
GATE 2 after the chapter map is proposed how the book is organized, before any prose is written
GATE 3 at the end that the book is finished: every bridge approved, every flag dealt with

The dashboard calls the written pieces stories; the printed book still has chapters, and the files keep their names (chapters/, data/chapters.csv).

6.1 How it fits together

transcripts + records ─► story units ─► chapter map ─► chapters ─► book PDF
                          content/units/  data/chapters.csv  chapters/   output/
                                              ▲                          ▲
                                           GATE 2                     GATE 3

Story units are worth cutting even if you never print a book: their people, places and subjects front matter is how the Familypedia knows which passages of which recording concern which article.

6.2 Story units

A story unit is one story, memory or explanation that stands on its own: typically one to five minutes of tape, stored as one file in content/units/. Each unit holds the exact clean transcript excerpt (## Source), its metadata (people, places, subjects, timeline events, part, era or relative), and later the finished prose (## Shaped) plus notes for you.

Why units instead of chapters? Chapter boundaries move. A life stage splits in two, or an uncle turns out to deserve his own chapter. Because chapters are generated from units, moving a story is a one-line metadata change and a rebuild, not a rewrite. All editing happens in units, and generated chapter files are never edited by hand.

Signs that a new unit is starting: a question that changes the subject, a jump in time or place, "and another time…", "that reminds me…". A story told in pieces across sessions is one unit with several spans. A story told twice is one unit: the fuller telling gets shaped, and the differences go in its notes.

content/boundaries.csv

You (or Claude) list the cut points. For the sample it would be:

session,start,id,title,part,section,tier,people,places,events,flags
S1,00:00:02,U001,The boots,family,Others in the Family,told,Anders Calder,"Norway;Duluth, Minnesota",E001;E002,
S1,00:00:29,U002,The ore dock,life,The Ore Dock,,,"Duluth, Minnesota",E003,
S1,00:00:48,U003,The big snow,life,The Ore Dock,,Pete,"Duluth, Minnesota",E004,
S2,00:00:03,U004,The shoes,life,Walt and the Cabin,,,"Minneapolis, Minnesota",E005,
S2,00:00:17,U005,The dance,life,Walt and the Cabin,,Walt,"Minneapolis, Minnesota",E006,
S2,00:00:36,U006,The cabin,life,Walt and the Cabin,,Walt,"Pike Lake, Minnesota",E007,

Then generate the unit files and check coverage:

python $LINEAGE/scripts/make_units.py --check    # coverage only
python $LINEAGE/scripts/make_units.py            # write content/units/U###-*.md

The coverage check

Because each cut runs to the next one, every paragraph lands in exactly one unit or in the excluded list. That makes loss checkable instead of a hope:

paragraphs: 14 in 2 session(s)
  in units:   14  (6 units)
  excluded:   0  (0 spans)
  UNCOVERED:  0

The script exits with an error if anything is uncovered. Two more checks:

python .claude/skills/content-separator/scripts/units.py coverage   # every subject paragraph accounted for
python .claude/skills/content-separator/scripts/units.py check      # missing fields, uncited or Q&A-style Shaped text

Re-running make_units.py after you change boundaries is safe for the fields it knows: chapter, order, lead_in, break_before, status, ## Shaped and ## Notes carry over. Frontmatter it doesn't know about is rewritten away, though. That includes lead_in_approved, kind, and any date_display you added by hand. An apparatus unit (section 6.5) has no spans, so it shows up as "not in boundaries.csv". Leave it in place and don't use --prune while you have one. New units start with break_before: false; set it to true where you want a ❧ break between stories.

6.3 The chapter map

Claude proposes the map (the chapter-index-builder skill) from the timeline and the units, not from the order things were said. It writes the map to data/chapters.csv.

Two parts: - The family part ("Those Who Came Before"), first by default: one chapter per named relative, titled with just the name, oldest generation first. Relatives with only a story or two share a chapter called "Others in the Family", each under their own heading. - The life part: the subject's own life, chronological, one stage or place per chapter. A theme that spans decades can be its own chapter, placed where it peaks.

Chapters are at least ten pages by default (chapters.min_pages in book.yaml, measured by building; make chapter tells you when one is short). A chapter gets there with real material: the subject's full stories in their own words, the records, built-out sourced context. Never with padding. Where the material is thin, chapters are combined (a relative with one story joins "Others in the Family"; two thin stages of life become one) rather than printed thin. Every chapter ends with THE RECORDS, listing every source found and used.

Sizing. Aim for roughly 1,000–2,500 words of the subject's speech per life chapter. Under about 600, merge with a neighbour; over about 3,000, split at a natural turn. A relative's chapter can be a single page.

Columns per chapter (data/chapters.csv):

column what goes in it sample
chapter, file, title, part order, file name, title, part name 2, chapters/02-the-ore-dock.typ, The Ore Dock, Her Life
setting, dates the place-and-years line under the title Duluth, Minnesota, 1938–1950
summary_line a short line in the old-book manner, never a list On Tin Pails and Tunnels
epigraph, epigraph_source the chapter's one epigraph: public-domain, verified (section 6.4)
columns 2 (two justified columns) by default; 1 for a chapter in one measure 2
summary 2–4 neutral sentences, for you and the introduction
status proposed → approved → drafted → reviewed → final approved

Titles are plain stage or place names, or a phrase the subject said.

GATE 2: approve the map

Claude shows you a simple outline (number, title, setting, years, one line each). Nothing is shaped or assembled until every row's status is past proposed. Shaping against the wrong structure is the most expensive mistake in the pipeline, so take your time here. Move stories around, merge, rename.

After approval, each unit gets a chapter and an order, and the chapter number is written back into the timeline. To see what lands where, and what landed nowhere:

python .claude/skills/chapter-generator/scripts/assemble.py --list

A unit that fits nowhere goes on your list. It is never silently dropped.

6.4 Writing

The stories are a biography written about the subject in the third person, past tense, by the family member who recorded it (you, the author). They should read like a real book, not an interview. The master rulebook is .claude/skills/memoir-style-guide/SKILL.md, and every other writing skill defers to it. Here are the rules that matter most.

The rules

Voice - Third person, past tense: "Ruth left Duluth in 1956." The subject never narrates. - The author never appears in a chapter. Your voice belongs in the introduction and an optional afterword only. No "my grandmother", no "I asked her". - No interview format anywhere. No questions, no "when asked", no "she recalled in an interview", no speaker labels. A question's content is folded into the answer as a plain statement, and both timestamps are cited. (If #asked ever appears in a chapter, the template refuses to compile.) - Tone: warm, plain, lightly wry about the world, never at the subject's expense.

Facts: narration may restate, order and connect. It may not add. - No feelings, thoughts or motives she didn't state. No weather, sensory detail or dialogue she didn't give. No "she must have…". - Hedges carry over. "I think it was '72" becomes "around 1972", never "in 1972". - A bare "yeah" to a leading question is a weak fact. Hedge it or flag it; don't state it. - Interpretation ("it was the end of her childhood") goes in a bridge for you to approve. - Unclear names, dates or relationships become questions in facts/gaps.md, never guesses. - Sensitive material (living people, legal trouble, health, violence, trauma) is written accurately and flagged with // REVIEW:. It is never cut or softened by the writer. You and the family decide what prints.

Quotation: how her voice survives - About one or two direct quotes per page. A chapter with no quotes has lost the subject. - Quotes are exact from transcript/clean/. Only fillers and false starts are removed, with "…" for an internal cut. Her grammar and dialect stay. - Quote the verdicts, the humor and the punchlines, and paraphrase the logistics. - Short quotes run inline. Passages of 40+ words go in #verbatim[...], at most a few per chapter. - Never invent or tidy a quote. If the wording isn't on tape, it's paraphrase.

Sourcing: every paragraph ends with a comment that never prints

// src: [S1 00:00:29]-[S1 00:00:47]        what the paragraph rests on
// src: transition                          pure connective text; no new facts allowed
// src: context                             a paragraph about the world, with...
// context: <the claim> — <source, page or URL>   ...one line per borrowed fact
// src: record R014                         a document from data/archives.csv
// REVIEW: <a doubt or sensitive item for the author>
// NOTE: <a research caveat>

Context about the world (what Duluth was like then, what a nurse earned) is welcome, but: it describes the world, never the subject. It stays local to the chapter's place and years. Every fact is sourced on a // context: line (a real source, never "general knowledge") and listed in the chapter's THE RECORDS, and it takes up no more than about a quarter of a chapter. Juxtaposition is allowed and causation is not: "That spring the mill cut its hours" is fine, but "so she left" is not unless she said so. For a real example, see section 7 of EXAMPLE-CHAPTER.md:

// context: Elmo P. Hohman, The American Whaleman (1928), p. 15 (green hand's lay 1/200),
// p. 240 (about 20¢ a day vs about 90¢ for unskilled labor ashore)

Epigraphs: each chapter opens with one (two allowed in a family-history chapter; none if you ask), from a writer of the chapter's place and time, public-domain literature only, never quoted from memory, and checked with verify_quotes.py against a saved copy of the text listed in facts/sources/works.csv.

Before and after: one story, three ways

This is the sample's unit U002. Here is the clean transcript:

**Grandma** [S1 00:00:29] When I was little, maybe seven or eight, I'd walk my dad's lunch
down to the ore dock. A tin pail. And the dock was so tall you had to tip your head all the
way back. The trains ran right out on top of it.

**Sam** [S1 00:00:44] Were you scared of it?

**Grandma** [S1 00:00:45] No. Well, yeah. A little.

WRONG: smoothed. This version reads nicely and is full of things she never said:

On warm summer mornings, seven-year-old Ruth proudly carried her father's lunch down to the ore dock. The towering structure terrified her: "you had to tip your head back to see the top." She loved those walks with all her heart.

QUOTE NOT FOUND chapters/99-wrong.typ:4: "you had to tip your head back to see the top."
    not in transcript: "you had to tip your head back to see the top"

WRONG: interview format. This version is accurate and still wrong:

When Sam asked whether the dock scared her, Ruth recalled: "No. Well, yeah. A little."

RIGHT. This is the sample's shaped text, exactly as it is in content/units/U002-the-ore-dock.md:

When she was seven or eight, as best she remembered, Ruth walked her father's lunch down to
the ore dock#idx("Duluth, Minnesota!ore dock") in a tin pail. The dock was so tall "you had to tip your head all the way back,"
and the trains ran right out on top of it. Whether it frightened her, she never quite
settled: "No. Well, yeah. A little."
// src: [S1 00:00:29]-[S1 00:00:47]

Two more patterns from the sample: - Folding a question. Sam asked "And that's where you met Grandpa?" and she answered "At a dance. A church dance, in 1959." The book says: "She met Walt at a church dance in 1959." (U005) - Keeping a hedge. "We built the cabin on Pike Lake ourselves, starting in, oh, 1964 or so" becomes "Starting in 1964 or so, Ruth and Walt built a cabin on Pike Lake with their own hands." (U006)

Bridges

A bridge is any sentence the writer added that isn't plain fact: an interpretation, or a connecting line that might color the story. In the sample, unit U005 has this lead-in:

lead_in: "Three years later, the thing she remembered best about Minneapolis had nothing to do with nursing."

Nothing on the tape says that, so it's assembled as #bridge[...]. In drafts it prints highlighted, like this: ⟦BRIDGE: …⟧. The final build refuses to compile while any bridge remains, and an unapproved bridge is never narrated. To decide on one: - Approve a lead-in: add lead_in_approved: true to that unit's frontmatter. It then prints as plain text. - Approve a bridge inside Shaped text: replace #bridge[...] with the sentence itself. - Reject it: delete it.

make status shows how many bridges are pending, and the Stories tab can filter to the stories that still have them.

Writing the units and checking quotes

Ask Claude to "shape the units for chapter 2" (content-separator plus the style guide). It writes each unit's ## Shaped section, usually with:

python $LINEAGE/scripts/shape.py U002 drafts/U002.md        # writes ## Shaped, status: shaped
python $LINEAGE/scripts/shape.py U002 --status approved     # only you set approved

Then check every quotation in the chapters against what the subject actually said:

python $LINEAGE/scripts/verify_quotes.py --transcript chapters/[0-8]*.typ

This matches every double-quoted passage of three or more words against the subject's paragraphs in transcript/clean/. Case and punctuation are ignored, and nothing else is. A quote found only in the interviewer's words is reported as "check attribution". A quotation from a letter or another relative is exempted by putting // quote-source: <where it's from> in the same paragraph. make draft runs this check and warns, and make final stops on any failure.

Before a unit or chapter is done, Claude runs the style guide's checklist. Read it yourself at least once (section 9 of the style guide). The scripts catch speaker labels, #asked, missing citations and misquotes. They do not catch "when asked" phrasing, invented feelings or firmed-up hedges. Those need a reader.

6.5 Family-history chapters

Stories about people nobody alive has met are the most fragile part of any record. Usually they're secondhand, sometimes contradictory, and once printed they become "what happened". The family-history-chapters skill adds rules on top of the style guide. Read EXAMPLE-CHAPTER.md alongside this section: it's a real ancestor's chapter, annotated rule by rule. The research behind such a chapter is section 4.

Tiers and the chain of telling

Every fact sits in exactly one tier: - witnessed: the subject saw it herself; - told: someone named told her ("as her father told it"); - lore: the family's story, source unclear ("that's what they always said").

A fact a record supports is documented, and its record is cited. The chain stays visible in the prose: "According to Ruth's father, Anders crossed at sixteen", not "Anders crossed at sixteen". Lore keeps its hedges.

The chapter's first paragraph says what kind of material this is and who it passed through, and the opening section makes a promise: "Where the records and the family part company, this chapter says so." The sample's chapter 1 does both in one paragraph (content/units/U001-the-boots.md).

Line of descent

A box at the foot of the opening page runs from the earliest known ancestor down to the living family, with marriages included. It's built only from documents and links you've confirmed, and unconfirmed links are marked. From the sample:

#descent(
  ([Anders Calder (link from Norway unconfirmed)], none),
  ([Ruth's father], none),
  ([Ruth Calder (b. 1938)], [Walt]),
  ([their children and grandchildren], none),
)

A small family tree

When you'd like a picture of the family around an ancestor rather than a single line, the chapter can carry a three-generation tree at the foot of its first page instead of (or beside) the line of descent: parents, the ancestor and spouse, their children, and the line continued under the right child. Same rules: documents and confirmed names only.

#family-tree(title: "The Family of Anders Calder",
  couple: ([Anders Calder (c.1875–?)], none),
  children: ([Ruth's father], [a sister (name not recorded)]), line: 0, descent: [and so to Ruth])

Anyone who served

For an ancestor who served in a war, the official service record comes first and is the spine, and the chapter is built out around it in time order: the town they grew up in and the country then; how the war drew their countrymen in, and those soldiers' reputation; how a young soldier enlisted, was examined, trained and shipped; what their unit met when they got there (its battles, weapons, generals); being sent home; and what veterans came home to. First-hand accounts from soldiers in or near the unit (diaries, letters, memoirs, unit histories) are quoted briefly and exactly, and saved so the quote check can verify them. A sensitive entry in a record (a punishment, a court case) stays out of the prose until you decide. Details: family-history-chapters skill, §6a.

THE RECORDS

Every family-history chapter, and any chapter with // context: lines, ends with a small-type THE RECORDS block after the closing paragraph. It lists where each documented claim came from: catalogue numbers, record titles, database IDs, newspaper titles and dates, books with years and pages, and links. It is drawn from the catalogue built in section 4.

#records(
  [*Census.* <census year, place, page and line>: Anders Calder, <age>, <occupation>.
   <where the free index is>. #link("https://…")],
  [*A note on the name.* <which of two similar records is meant, and why>.],
)

The angle-bracket parts are yours to fill in. The sample hasn't consulted any records yet, and its #records block says exactly that. Include a note on the name whenever a record could be confused with a similar one (section 4.4).

THE RECORDS and the photographs page (section 6.6) go in a final apparatus unit for the chapter (kind: apparatus, same section, highest order, no spans), so they survive reassembly. The sample keeps it inside U001 instead, which works for a one-unit chapter.

When documents contradict the family

Print both, and say so. Tell the life once, in order, with the record as the spine. Where the family's version differs, tell it at that point, attributed ("The family remembered his war differently…"). Let the record confirm what it can. Dates in narration follow the record, with the family's date given as theirs. Log the conflict in facts/gaps.md and mark it // REVIEW:. Section 4 of EXAMPLE-CHAPTER.md shows this done well: the family's "surgeon, drafted" set beside the records' "soldier, enlisted", and the bounty money that turned out to be real.

Where the Records Are

make draft turns the printable rows of data/archives.csv (section 4.1) into the back-matter appendix "Where the Records Are". Only rows with print_permission: yes print, and private details never do.

6.6 Photos and illustrations in stories

Images are catalogued in section 3.3. Putting them on a page adds these rules.

python .claude/skills/photo-processor/scripts/prepare_print.py --color bw
python .claude/skills/photo-processor/scripts/prepare_print.py --only P007 --crop P007=40,30,1880,1400

Allowed: rotating, cropping away the scanner bed or album page, converting to grayscale. Not allowed on a real photograph without the owner's say: colorizing, upscaling, face restoration, or removing people or objects. These fabricate detail. Flag the damage and let the owner decide whether to get a better scan or a professional restorer.

Resolution: 300 ppi at the printed size; 200 ppi is a hard floor. Below that, print it smaller or get a better scan. Never upscale to hide it. Note that make final's preflight measures resolution only for images placed with #photo(...). For #plate(...) images, check max_print_width_in in the photo index yourself.

Placing and captioning

Use few images, and only where they belong: a portrait near where a key person is introduced, and one or two pictures at the exact moments they show. Nothing decorative, and no more than one per page. Each image is anchored after the paragraph it belongs to, in the unit's Shaped text:

#plate("/photos/print/P014.jpg", caption: "Anders Calder", id: "P014")

About 3.9 in wide for portraits and 4.4 in for landscapes. #plate-pair sets two side by side, and span: false keeps a plate in one column of a two-column chapter.

Caption conventions:

Situation Caption
a portrait the name exactly as you give it: Anders Calder
a scene a short title in the text's own words: The cabin on Pike Lake
date and place confirmed Ruth and Walt at Pike Lake, 1966. Add a date or place only if confirmed.
identification uncertain Probably Pete, about 1950 or Believed to be the Duluth ore dock, or no caption until you ask
an illustration short like any other: The tunnel to the street. It is marked as an illustration in the photo index, the chapter note and the copyright page (below)
a map Positions from the logbook; lines between them are approximate.

If you supply a caption, it's used word for word, and any mismatch with the text gets a // REVIEW:. Images with no matching story go on a candidates list, not into a random chapter.

Illustrations (AI-generated or an artist's): strict rules

Often no photograph of an ancestor exists. A rendering is allowed only if all four hold:

Also: nothing that poses as an archival document, no caricature, nothing graphic. A real photograph may serve as the likeness reference (say so under (b)), but the photograph itself is never altered. To make illustrations in a consistent period style, with a relative's likeness locked from their real photographs, see "Generating illustrations" below.

Maps: drawn from data, never generated

A route CSV (section 4.3) becomes a printed map:

python $LINEAGE/scripts/make_route_map.py facts/records/anders/crossing_track.csv \
    --out photos/print/P050.png --title "The crossing" --subject "Anders" \
    --caption "Positions from the records; lines between them are approximate."

The --subject flag names whose journey it is in the legend. Without it, the legend uses the book's subject. The map is black and white and evidence-coded. Filled dots are places recorded with the subject aboard. Open dots are places recorded without them. A dashed line marks the subject's track, approximate between recorded points. Legs no record covers are not drawn and are labelled "not recorded". In the example, gap_before=yes on New York leaves the Atlantic crossing blank and labels it that way. The output is a 300 dpi PNG. Coastlines come from Natural Earth (public domain) and are downloaded once. For the map lettering to be EB Garamond, copy $LINEAGE/fonts/*.otf into your system font folder (~/Library/Fonts on a Mac, ~/.local/share/fonts on Linux). make_route_map.py --help lists the label and extent options.

The photographs page and permissions

At the end of each family-history chapter, after THE RECORDS, the photographs page (#photo-addendum) lists real photographs held by archives: catalogue number, date, description, link, and a note on what permission printing them would need. Thumbnails print in drafts, and in the final book only once the holder has given permission (show-images: true).

#photo-addendum(note: [Held by <the archive>; reproduction needs its written permission.],
  (none, [<catalogue number> · <date>], [<what the photograph shows>], "https://…"),
)

The first item in each row is a thumbnail path, or none until permission arrives.

Family photographs need permission too: ask whoever holds the original before it prints, and record it in the index (print_permission).

Generating illustrations

Plan every illustration in photos/images-plan.yaml (example in $LINEAGE/templates/images/): a period style preset (how photographs of that time looked), the scene's citation, the prompt, and optionally a likeness set: real photographs of a relative in photos/reference/<name>/, used as references so the person looks like themselves across a set. Then:

python $LINEAGE/scripts/generate_images.py --dry-run     # read every prompt first
python $LINEAGE/scripts/generate_images.py               # photos/generated/<id>.png + grayscale print copy

The key comes from OPENAI_API_KEY for the one command or from ~/.openai_api_key (chmod 600, outside every repo), and is never written into the project. Index each result as kind: illustration before placing it.

An illustrated story

For one scene the subject told in vivid detail, the chapter can carry an eight-panel sequence over two facing pages, each panel captioned with the subject's exact words:

#story-page(title: "The Night of the Storm",
  story-panel("/photos/generated/print/IMG-03-1.jpg", 1)[“The wind took the barn roof clean off.”],
  …)
// src: panel 1 [S2 00:14:05]; …
#story-page(note: [Illustrations after Ruth's account; artist's renderings, not photographs.], …)

6.7 Assembling and building

Generate the chapters

Once units are shaped, ask Claude to "assemble the chapters", or run:

python .claude/skills/chapter-generator/scripts/assemble.py        # all chapters
python .claude/skills/chapter-generator/scripts/assemble.py 2 3    # just these

Each chapter opens with its title, the setting · dates line, the summary line and any epigraph. The units follow in order, with ❧ between them, a drop cap on the first paragraph, and == headings where one chapter holds several relatives. Claude then does a seam pass. It reads each chapter start to finish and fixes jumps and repeated setup in the units, never in chapters/*.typ. Those files carry a GENERATED header and are overwritten on every build.

The Stories tab shows the same stories oldest first in era bands, with Read (real book pages rendered through the Typst template) on each row, story states (draft · in the book · kept aside), and marks on stories that are stale because a source changed. "Generate" there hands the work to the Genealogist in the terminal drawer.

Build a draft

make -C "$LINEAGE" draft PROJECT="$PWD"

In order, this runs: 1. sync: refreshes book/template.typ from the repo (the design lives in one file, $LINEAGE/book/template.typ); 2. assemble: regenerates every chapter from its units; 3. quote check: verify_quotes.py --transcript, which warns here and blocks in final; 4. appendices: "A Timeline" from facts/timeline.csv and "Where the Records Are" from data/archives.csv; 5. front, glossary, sources: the title and copyright pages from book.yaml, the reader's glossary from data/reader_glossary.csv, and "About the Recordings" from transcript/sessions.csv; 6. main: generates book/main.typ from data/chapters.csv; 7. compile: writes output/book-draft.pdf, padded to an even page count; 8. back index: reads the resolved index page numbers into data/back_index.csv.

Drafts show bridges highlighted, editor notes in red and image IDs. Afterwards, actually look at the title page, the contents, a part page, two chapter openers, a spread with a plate, and the index. Fix anything wrong at its source (a unit, chapters.csv, book.yaml), never in a generated file.

To look at one chapter without building the whole book, run build_book.py chapter chapters/NN-x.typ (→ output/preview-NN-x.pdf); add --final --pages for one image per page in output/pages/ to share outside the PDF.

Front matter

In print order: title, copyright, dedication, contents, then the introduction.

The index

Index terms are marked in the text as it's written, at first mention per paragraph, and they're invisible in print:

Her grandfather Anders#idx("Calder, Anders") worked the ore boats.#idx("ore boats")
#idx("Duluth, Minnesota!ore dock")            // a sub-entry
#idx-see("Grandpa Walt", "Calder, Walt")      // a cross-reference

People are entered as "Surname, Given", and places as "Town, State". Page numbers are computed during typesetting, so they stay right when pages move. After a build, review data/back_index.csv for duplicate spellings and entries with too many page references. For a spreadsheet of chapters and index:

python .claude/skills/chapter-index-builder/scripts/index_tools.py wordcount
python .claude/skills/chapter-index-builder/scripts/index_tools.py xlsx      # output/chapters.xlsx

Versions

Every PDF that leaves the project should be traceable to the exact text that produced it:

python .claude/skills/book-generator/scripts/version.py init      # once: git repo, originals' checksums
python .claude/skills/book-generator/scripts/version.py snapshot "chapter map approved"
python .claude/skills/book-generator/scripts/version.py snapshot "review copy for Ruth" --major --note "mailed Nov 3"
python .claude/skills/book-generator/scripts/version.py list

Take a snapshot at every gate, and always a major version before anything is sent to anyone. "Page 41, line 3" only means something against the exact copy they received. If you put the project on GitHub, keep the repository private, because transcripts are in it.

GATE 3: final build and preflight

make -C "$LINEAGE" final PROJECT="$PWD"

This builds the draft, re-checks every quotation (and stops on a mismatch), and compiles output/book-final.pdf in final mode. Final mode refuses to compile while any unapproved bridge remains, so the sample stops here on purpose:

error: panicked with: Unapproved #bridge left in manuscript: [Three years later, the thing she remembered best about Minneapolis had nothing to do with nursing.]

Then it runs the preflight for the trim, printer and color set in book.yaml: page size, even page count, the printer's minimum page count, embedded fonts, image resolution, and no leftover draft markers. A clean result ends with No blocking problems found.

Before you call it final: every bridge decided, every // REVIEW: read (grep -rn "REVIEW" content chapters), every question in facts/gaps.md answered or consciously left open. Then you, the author, sign off.

Where work stops, and the full first pass

By default work stops for you at three gates (speakers, the chapter map, the final build); between them it runs straight through and saves its questions for the next gate. You choose where it stops: add gates in book.yaml (workflow.gates) or in your project's CLAUDE.md ("one chapter at a time"), and the stricter rule wins.

When you'd rather see the whole book at once, ask for a full first pass. Claude runs the book swarm (BOOK-SWARM.md): for every chapter, two researchers, a writer, three editors (fidelity, your taste, layout), two independent judges with revisions, and a 100% sources pass; then front matter, an images plan, the index, the assembled draft and a whole-book continuity review. It works on a branch of your project, and nothing is final until you've read it.

6.8 Reviewing with the subject across a distance

The best check on the record is the person it's about. If they live far away, or don't use a computer, here's a round trip that works by mail and phone, and that brings back a new source as well as corrections. The repo has no dedicated tool for this. What follows is a small Typst file you add yourself, plus the normal pipeline.

1. Put your questions in the text

Where you need the subject's answer, add a draft-only note to the unit's Shaped text, right where the question arises:

"Nobody believes that," she said. #note[Question for Ruth: which Pike Lake was it? Minnesota has several.]

#note[...] prints in drafts only and never in the final book. Good questions come from facts/gaps.md, low-confidence timeline rows, bare-"yeah" facts, and // REVIEW: items you want her view on. (Home's Request more builds question lists from the same places.) Questions about other living people need more care.

2. Make a large-type review edition

Save this as book/review.typ in the project. It sets 14-point body type on US Letter paper (so it prints at home), numbers the lines on every page, leaves a wide margin for writing, and keeps the bridges and your questions visible. List your own chapter files at the end:

// Large-type review edition: 14 pt body, numbered lines, a wide margin for notes,
// and draft notes, bridges and questions printed. Not for the printer.
#import "/book/template.typ": *

#show: book.with(title: "The Lake Was Always There", trim: "8.5x11")
#set page(margin: (inside: 1in, outside: 2in, top: 1in, bottom: 1in))
#show par: set text(size: 14pt)
#show regex("⟦[^⟧]*⟧"): set text(size: 12pt)
#set par.line(numbering: "1", numbering-scope: "page", number-clearance: 1.2em)

#show: main-matter
#include "/chapters/01-others-in-the-family.typ"
#include "/chapters/02-the-ore-dock.typ"
#include "/chapters/03-walt-and-the-cabin.typ"

Build the draft first, so the chapters and template are current. Then:

make -C "$LINEAGE" draft PROJECT="$PWD"
typst compile --root . --font-path "$LINEAGE/fonts" book/review.typ output/review.pdf

Line numbers restart on each page, so a note can say "page 4, line 12". Two-column chapters stay two columns. Typst line numbering needs Typst 0.12 or newer.

3. Print it, send it, call

4. Feed the corrections back in as a new source

The call is a new recording, and it goes through intake and transcription like any other. Never edit the old transcripts.

cp ~/Downloads/review-call.m4a audio/
$LINEAGE/scripts/transcribe.sh         # only the new session is processed (say S3)

Earlier sessions are skipped. transcription_locked: true stops nothing here, because this session has never been transcribed. Confirm its speakers (GATE 1 again, for S3 only), render, and then:

6.9 Printing

Trim size (print.trim): 7×10 in is the default and suits a book with photographs. 6×9 works for a mostly text book. 8×10 and 8.5×11 are also supported. A black-and-white interior (print.color: bw) costs several times less on print-on-demand than color, and red details print as gray. Trim and printer can also be set in Settings → Project settings.

Minimum page counts, as the preflight checks them:

Printer (print.printer) Minimum pages
kdp (Amazon KDP paperback) 24 (hardcover: 75)
ingramspark 18
lulu 32
blurb 20

The build always pads to an even page count. Front and back matter add up quickly. The sample, from two minutes of tape, is 30 pages, which is enough for KDP, IngramSpark and Blurb, but not for Lulu.

What Lineage doesn't do: - The cover. Each printer has a cover calculator that sizes the spine from your final page count and paper. Make the cover once the page count is final. - PDF/X conversion. IngramSpark prefers PDF/X-1a or X-3. If it rejects the file, convert with Ghostscript or Acrobat. - Full bleed. Bleed is only needed for images running off the page edge, and the template doesn't do it.

Proofs. Order a printed proof before anything else, and read it on paper. Pages that looked fine on screen show their problems in print: a photo too dark, a widow, a caption that crowds its image. Snapshot each proof as a major version so your notes match the copy. Allow a week or two per proof round for printing and shipping, and usually plan on two rounds. If the book is for a particular birthday or reunion, order the first proof at least a month before.

6.10 Narration

The Stories tab can narrate any story with ElevenLabs (add the key under Settings → Connectors). Each row has Listen beside Read, and Narrate / Re-narrate; each story can have its own voice (the default is set in Project settings). Narrate all shows the character count before it spends anything. Audio is marked stale when its story changes, plays in a pinned player, and can be downloaded.

The narration script is extracted from the story itself, and an unapproved bridge is never narrated: decide the bridges first (section 6.4).

Narration is written but has not been tested end to end with a real key; treat the first run as a trial. The story's text goes to ElevenLabs to be read aloud.


7. Troubleshooting

UnpicklingError / WeightsUnpickler / "weights_only" when WhisperX starts. PyTorch 2.6 and later refuse the speech-detection checkpoints WhisperX loads. Use $LINEAGE/scripts/wx.py in place of the whisperx command, as transcribe.sh does. It applies the narrow fix in scripts/hf_compat.py.

TypeError: ... unexpected keyword argument 'use_auth_token'. huggingface_hub 1.0 removed that keyword, and pyannote 3.4 still passes it. hf_compat.py renames it. Run diarization through diarize.py, diarize_chunked.py or transcribe.sh, which all import the shim. Don't fix either error by pinning older torch or huggingface_hub, because that breaks WhisperX. The tested versions are in requirements-transcribe.txt.

pip can't find torch==2.8.0. Your Python is probably newer than the pinned PyTorch supports. The stack was last run on Python 3.13. Rebuild the venv with an older Python: rm -rf "$LINEAGE/.venv" && python3.13 -m venv "$LINEAGE/.venv", then make install install-transcribe.

401, 403 or GatedRepoError when diarization starts. The token is wrong, or the licence wasn't accepted on both pyannote pages (speaker-diarization-3.1 and segmentation-3.0) while logged in as the account that owns the token. See section 1.1. Your transcription is safe. Fix the token and run python $LINEAGE/scripts/diarize.py.

Diarization runs for hours. One long session was labelled in one piece. Make sure you're running through diarize.py, which chunks anything over 20 minutes by default, rather than transcribe.sh --one-pass, which doesn't chunk. Or force chunking with python $LINEAGE/scripts/diarize_chunked.py S4.

Names in the transcript as fake speech. This is the initial prompt being echoed. Shorten proper_nouns, then cut the echoes with drop_segments/scrub_inline in corrections.json, and re-render. Don't re-transcribe a cited session.

A scan or image-only PDF shows no text in Sources. OCR needs tesseract (brew install tesseract / apt install tesseract-ocr). Install it and re-ingest the source, or paste the text in the edit drawer.

Lineage vX.Y.Z is not installed from make dashboard. The release pinned in lineage.lock isn't under ~/.lineage/releases/ on this machine. Run the command it prints: make update-lineage TO=vX.Y.Z APPLY=1.

unknown font family warnings, or the book comes out in the wrong typeface. You compiled without the bundled fonts. Always pass --font-path "$LINEAGE/fonts" to typst compile, or build through make draft, which does it for you. Check that $LINEAGE/fonts/ contains the EB Garamond .otf files.

failed to download package @preview/droplet. The first build needs internet access once to fetch the drop-cap package. After that it's cached.

The final build stops on Unapproved #bridge left in manuscript. This is working as intended. The message names the chapter and line. Approve the bridge (lead_in_approved: true on the unit, or replace #bridge[...] with the sentence in the Shaped text) or delete it, then build again. make status shows how many remain.

The final build stops on QUOTE NOT FOUND. A quotation doesn't match the transcript word for word. Fix the quote to match transcript/clean/ or turn it into paraphrase. If it's someone else's words, add // quote-source: <where> to the paragraph.

Odd page count in the preflight. build_book.py compile pads to an even count by recompiling with a blank last page. If you compiled book/main.typ by hand, rebuild through make draft or make final.

typst not found. Install Typst (brew install typst, or see https://github.com/typst/typst#installation) and check that typst --version works in the same terminal you run make from.

pdfinfo or pdffonts not found, or the preflight crashes. Install poppler (brew install poppler / apt install poppler-utils). make install calls it optional, but the preflight in make final needs it.

whisperx does not import. Run make install-transcribe. The transcription stack goes into the same venv as everything else.

A folder named ~ appeared inside the repo. PROJECT=~/… wasn't expanded by your shell (zsh). Delete that folder and use PROJECT="$HOME/…".

A chapter edit vanished. You edited a GENERATED file in chapters/. Make the change in the unit or in data/chapters.csv and rebuild.


8. Honest limits

What needs a human

The tools will not, and should not, decide these for you: - Speaker confirmation (GATE 1): which voice is the subject. - Accepting tag suggestions in the Familypedia, and applying a genealogy rebuild after reading what it changed. - Reviewing contributors' uploads before they join the record. - Identifying people in photographs, and permission to print them. - Every question in facts/gaps.md, and every contradiction: the platform shows both sides, and only a person can weigh them. - For a book: approving the chapter map (GATE 2), approving or rejecting every bridge (meaning every interpretation), every // REVIEW: item, and final sign-off (GATE 3). - Sensitive material: what to print or share about living people, illness, legal trouble, family conflict. The writer is told to flag it and never cut or soften it on its own, so the decision is yours and the family's.

The scripts check what can be checked mechanically: quotes against the transcript, coverage, citations, missing fields, bridges, evidence on genealogy links, page counts. They can't tell whether a sentence quietly invented a feeling. Read what it writes.

What isn't built yet

The dashboard's app/ROADMAP.md is the current list. Among the things this guide does not describe as working: the contributor-facing page behind an invite link (waits for hosting); a photo gallery and lightbox with annotation and people tagging on the image; a screen for pasting and scanning public-record links into the catalogue; the impact pass and "update everything this affects" after new material arrives (Home's Needs you already surfaces stale stories); and stories with hyperlinked people, places and citations and images editable in place. Google Drive sign-in and sync, crest generation and ElevenLabs narration are written but not tested end to end.

How long it really takes

Machine time is the easy part to estimate. Speech recognition with large-v3 on a laptop CPU runs at roughly real time or slower, so an hour of tape takes an hour or more. A CUDA GPU is much faster. Chunked speaker labelling costs about half a minute of compute per minute of audio on a CPU. You can leave both running overnight.

Your time is the larger part, and it doesn't shrink much with better tools: - listening to confirm speakers and checking a few random stretches against the audio; - reviewing tags, genealogy rebuilds and contributors' uploads; - research into records, which can take longer than everything else combined; - photographs: finding them, getting them scanned, asking who's in them; - for a book: reading and approving the chapter map, reading every chapter against the tape, deciding every bridge and // REVIEW: item, the review round trip by mail (weeks), and one or two printer proofs (a week or two each).

A short book from a few hours of tape is a project of weeks of evenings, not a weekend. The record itself is never finished: it grows as the family adds to it.

How many pages your recordings will make

This is a rough estimate. Your mileage will vary with how much your subject talks, how much research you add, and how many photographs you have.

Taken together: plan on roughly 10 pages of narrative per recorded hour. Add the openers, photographs, context and back matter on top of that. Chapters with a lot of research grow well beyond it: in the worked example, about 25 minutes of tape plus public records became a 10-page two-column chapter. No measured figure exists yet for a whole book. For your own number, build a draft after the first few chapters and scale up from that.

Where your material goes

Originals stay on your computer, unless you turn on the Google Drive connector, which copies sources to your own Drive folder. Speech recognition and speaker labelling run locally. The Hugging Face token only downloads the model. The dashboard binds to 127.0.0.1 only.

Text goes to cloud models. Claude Code does the model work, so transcript passages, extracted text, notes and drafts are sent to Anthropic while you work. With an Anthropic key set in Connectors, the Sources tab's understanding pass sends a source's extracted text (and a small image, for its description) to Anthropic too. Narration sends a story's text to ElevenLabs. If a recording or document contains something that must not leave your machine, keep it out of the project, or cut it from the transcript (exclude the span in boundaries.csv, or leave that session out of the writing stages) before you start. Research requests go to the sites concerned, and never with your personal details. See PRIVACY.md.

The record is yours, in plain files. Everything Lineage builds is CSV, JSON, Markdown and Typst in the project folder, with raw snapshots of the records it used. Keep the project in a private git repository and it outlives the links, the apps and the people in it.