Hear the missing layer
Show how identical words can become different ads through pace, emphasis, breath, and emotional posture.
PUBLIC BUILD / 2026
The video is strongest as a transparent build story: a real adapter, a corrected misunderstanding, an unfinished evaluation, and a clearer experiment at the end.
A YouTube and X narrative built around process rather than a victory lap.
01 / Story
The viewer should leave understanding what was trained, what evidence exists, why the renderer changed, and what would count as a real win. The tension is not whether the code ran. It is whether a small supervised dataset changed useful behavior.
I trained Inkling on the delivery of 42 real UGC ads. The adapter exists and the pipeline works. But I do not yet have enough evidence to tell you the voice direction is better.Claim-safe cold open
Show how identical words can become different ads through pace, emphasis, breath, and emotional posture.
Select performances, label delivery, quarantine bad rows, and train the adapter.
Explain the weak evaluation, the renderer decision, and the better next test.
02 / YouTube
“I trained an audio-language model to turn how winning UGC ads sound into performance direction. The checkpoint is real. The result is not proven.” Show the four-stage pipeline and the pending evaluation badge.
Demonstrate how a flat transcript loses the smile, rush, pause, hesitation, or certainty that made the delivery persuasive.
Clarify that Inkling reads audio and text, outputs text, and does not synthesize the final voice. Introduce Fish as the downstream renderer.
Show a sanitized version of the local selection interface: 3,000+ source ads, 50 chosen, no private media or client identity on screen.
Use a synthetic example to show source words, global direction, segment cues, and the word-preservation gate.
Explain the one-row proof, then the 42-row cohort: 31 text-only, 11 audio-conditioned, two rank-32 passes, 42 optimizer steps each.
Show the incident honestly: connection failure before step zero, dependency correction, then the same bounded run succeeding.
Explain existing script → same script with expressive direction. Separate that from the broader “write a new ad” prompt.
Describe the closed-list project policy, the discovery that it was too restrictive, and Fish's natural-language expressive cues.
Show Generate Ads as exploratory and Compare Markup as controlled. Then show the same-voice blind A/B plan—without inventing a winner.
End with 100–300 reviewed pairs, 20–30 untouched scripts, separate writing and direction stages, and a declared listening gate.
03 / X cut
“I had 3,000 high-performing UGC ads—and I wanted to train the part normal transcripts erase.” Visual: waveform becomes expressive text.
“I selected 50 performances, kept 42 valid labels, and trained a rank-32 Inkling adapter in two passes.” Visual: 50 → 42 → 84 steps.
“Inkling does not make the voice. It writes performance direction. Fish will render both models with the same voice.” Visual: Inkling text → Fish audio.
“The adapter exists, but I trained mostly script-to-markup—not product-brief-to-great-ad. So a flashy Red Bull prompt is not the clean test.”
“The real test locks one script, preserves every word, renders both outputs blind, and reveals identity only after I choose.”
“Pipeline proven. Improvement pending. That is the result—and the next experiment is better because of it.”
04 / Safety
Use the public site, synthetic examples, green test summaries, abstract model diagrams, and redacted dashboards. Do not improvise a terminal tour while authenticated provider state is visible.
| Safe to show | Keep off-screen |
|---|---|
| Aggregate counts, task structure, sanitized UI, test results | Environment files, API values, request headers |
| Generic base → adapter → renderer diagrams | Private checkpoint, session, or sampler identifiers |
| Synthetic scripts and expressive examples | Raw ads, client names, source URLs, private transcripts |
| “Checkpoint saved” and step totals | Local machine paths, hashes, provider logs |
| Pending comparison interface | A fabricated or prematurely revealed A/B winner |
Title: “I Trained an Audio Model on 42 UGC Ads—Did It Learn the Delivery?”
Alternative: “Can AI Hear What Makes a UGC Ad Work?”
Thumbnail: one waveform on the left, one expressive script on the right, and the restrained line “PIPELINE ≠ PROOF.” Avoid model logos and exaggerated result claims.
This is an independent experiment using a private ad library and hosted training services. Source media, credentials, and checkpoint identities are not public. Training completion is verified; expressive-quality improvement remains pending a blind listening evaluation.