INKLINGUGC PERFORMANCE LAB

PUBLIC BUILD / 2026

08 / Next

Proposed / not yet trained

A better training experiment

The next run should be larger, narrower, and easier to falsify: reviewed script-to-markup pairs, a locked holdout, and blind audio judgments with one renderer voice.

Reviewed script-to-markup pairs, untouched evaluation scripts, and blind listening.

01 / Dataset v2

Build 100–300 reviewed expressive pairs

The next training set should contain 100–300 reviewed examples of one narrow transformation: approved script in, Fish-compatible expressive markup out. Every target should preserve spoken words and pass a human review for coherence, restraint, and usefulness.

SourceApproved script

No ambiguous transcript or hidden copy edit.

DraftAudio-grounded direction

Use source performance where rights and quality allow it.

ReviewHuman-corrected target

Keep cues justified, coherent, and renderer-aware.

FreezeVersioned pair

Bind the exact input and target before training.

Quantity alone is not the goal. One hundred diverse, carefully reviewed pairs are more informative than thousands of noisy auto-labels that conflate writing, transcription, and performance.

02 / Holdout

Keep 20–30 untouched scripts out of training

Reserve 20–30 untouched scripts before teacher labeling or fine-tuning. They should span hook styles, emotional arcs, pacing demands, and product categories without duplicating creators or near-identical copy from training.

Coverage

Hook diversity

Confessional, contrarian, demonstrative, problem-led, and direct-response openings.

Coverage

Energy diversity

Quiet credibility, playful surprise, urgency, reassurance, and controlled excitement.

Leakage gate

Creator and copy separation

Near-duplicates cannot make a holdout look easier than it is.

The holdout is frozen once. If it is repeatedly used to revise prompts or labels, it becomes a development set and a new final test must be created.

03 / Separate skills

Train the writer and director as different stages

A practical production chain can still generate a complete ad. It should do so explicitly: first create or approve the copy, then direct its performance. Each stage gets its own prompt contract, data, and evaluation.

Stage A

Creative writer

Product facts and audience insight become a claim-safe UGC script. Evaluation judges hook, structure, specificity, and factual discipline.

Stage B

Expressive director

The locked script becomes renderer-ready performance text. Evaluation judges wording preservation, cue quality, and heard result.

This separation also reveals which model deserves fine-tuning. A strong general writer may need no adapter, while the expressive director may benefit substantially from domain-specific examples.

04 / Release gate

Define success before listening

Every holdout script produces a matched base-versus-adapter pair. Both texts are rendered by Fish with the same voice and settings. Reviewers listen blind, record A, B, tie, or neither, then reveal identity only after the decision is locked.

GateMinimum evidenceWhy
CoverageEvery frozen holdout has both valid armsMissing failures cannot disappear from the denominator.
WordingControlled scripts preserve spoken wordsThe test remains about direction.
ListeningAt least 90% of pairs receive a complete blind listenText appearance is not substituted for audio judgment.
PreferenceThe trained arm wins a declared threshold of non-tied decisionsSuccess is decided before identities are known.
Qualitative reviewFailures are categorized, not averaged awayMissed pauses and over-direction suggest different fixes.
Decision rule

If several holdouts need material correction, improve the labels or objective before buying more training. A larger run should follow better evidence, not replace it.