Hook diversity
Confessional, contrarian, demonstrative, problem-led, and direct-response openings.
PUBLIC BUILD / 2026
The next run should be larger, narrower, and easier to falsify: reviewed script-to-markup pairs, a locked holdout, and blind audio judgments with one renderer voice.
Reviewed script-to-markup pairs, untouched evaluation scripts, and blind listening.
01 / Dataset v2
The next training set should contain 100–300 reviewed examples of one narrow transformation: approved script in, Fish-compatible expressive markup out. Every target should preserve spoken words and pass a human review for coherence, restraint, and usefulness.
No ambiguous transcript or hidden copy edit.
Use source performance where rights and quality allow it.
Keep cues justified, coherent, and renderer-aware.
Bind the exact input and target before training.
Quantity alone is not the goal. One hundred diverse, carefully reviewed pairs are more informative than thousands of noisy auto-labels that conflate writing, transcription, and performance.
02 / Holdout
Reserve 20–30 untouched scripts before teacher labeling or fine-tuning. They should span hook styles, emotional arcs, pacing demands, and product categories without duplicating creators or near-identical copy from training.
Confessional, contrarian, demonstrative, problem-led, and direct-response openings.
Quiet credibility, playful surprise, urgency, reassurance, and controlled excitement.
Near-duplicates cannot make a holdout look easier than it is.
The holdout is frozen once. If it is repeatedly used to revise prompts or labels, it becomes a development set and a new final test must be created.
03 / Separate skills
A practical production chain can still generate a complete ad. It should do so explicitly: first create or approve the copy, then direct its performance. Each stage gets its own prompt contract, data, and evaluation.
Product facts and audience insight become a claim-safe UGC script. Evaluation judges hook, structure, specificity, and factual discipline.
The locked script becomes renderer-ready performance text. Evaluation judges wording preservation, cue quality, and heard result.
This separation also reveals which model deserves fine-tuning. A strong general writer may need no adapter, while the expressive director may benefit substantially from domain-specific examples.
04 / Release gate
Every holdout script produces a matched base-versus-adapter pair. Both texts are rendered by Fish with the same voice and settings. Reviewers listen blind, record A, B, tie, or neither, then reveal identity only after the decision is locked.
| Gate | Minimum evidence | Why |
|---|---|---|
| Coverage | Every frozen holdout has both valid arms | Missing failures cannot disappear from the denominator. |
| Wording | Controlled scripts preserve spoken words | The test remains about direction. |
| Listening | At least 90% of pairs receive a complete blind listen | Text appearance is not substituted for audio judgment. |
| Preference | The trained arm wins a declared threshold of non-tied decisions | Success is decided before identities are known. |
| Qualitative review | Failures are categorized, not averaged away | Missed pauses and over-direction suggest different fixes. |
If several holdouts need material correction, improve the labels or objective before buying more training. A larger run should follow better evidence, not replace it.