INKLINGUGC PERFORMANCE LAB

PUBLIC BUILD / 2026

09 / Video Guide

Production guide ready

Film the evidence, including the failure

The video is strongest as a transparent build story: a real adapter, a corrected misunderstanding, an unfinished evaluation, and a clearer experiment at the end.

A YouTube and X narrative built around process rather than a victory lap.

01 / Story

The experiment, not the victory lap

The viewer should leave understanding what was trained, what evidence exists, why the renderer changed, and what would count as a real win. The tension is not whether the code ran. It is whether a small supervised dataset changed useful behavior.

I trained Inkling on the delivery of 42 real UGC ads. The adapter exists and the pipeline works. But I do not yet have enough evidence to tell you the voice direction is better.
Claim-safe cold open
Act I

Hear the missing layer

Show how identical words can become different ads through pace, emphasis, breath, and emotional posture.

Act II

Build the training pipe

Select performances, label delivery, quarantine bad rows, and train the adapter.

Act III

Refuse easy certainty

Explain the weak evaluation, the renderer decision, and the better next test.

02 / YouTube

A 10–12 minute narrative spine

Cold open

“I trained an audio-language model to turn how winning UGC ads sound into performance direction. The checkpoint is real. The result is not proven.” Show the four-stage pipeline and the pending evaluation badge.

The missing data

Demonstrate how a flat transcript loses the smile, rush, pause, hesitation, or certainty that made the delivery persuasive.

What Inkling does

Clarify that Inkling reads audio and text, outputs text, and does not synthesize the final voice. Introduce Fish as the downstream renderer.

Select the cohort

Show a sanitized version of the local selection interface: 3,000+ source ads, 50 chosen, no private media or client identity on screen.

Create expressive targets

Use a synthetic example to show source words, global direction, segment cues, and the word-preservation gate.

Train carefully

Explain the one-row proof, then the 42-row cohort: 31 text-only, 11 audio-conditioned, two rank-32 passes, 42 optimizer steps each.

The failed attempt

Show the incident honestly: connection failure before step zero, dependency correction, then the same bounded run succeeding.

What the adapter learned

Explain existing script → same script with expressive direction. Separate that from the broader “write a new ad” prompt.

Why ElevenLabs became Fish

Describe the closed-list project policy, the discovery that it was too restrictive, and Fish's natural-language expressive cues.

Compare without cheating

Show Generate Ads as exploratory and Compare Markup as controlled. Then show the same-voice blind A/B plan—without inventing a winner.

The better next run

End with 100–300 reviewed pairs, 20–30 untouched scripts, separate writing and direction stages, and a declared listening gate.

03 / X cut

A 75–90 second version

00–08s

Hook

“I had 3,000 high-performing UGC ads—and I wanted to train the part normal transcripts erase.” Visual: waveform becomes expressive text.

08–22s

Method

“I selected 50 performances, kept 42 valid labels, and trained a rank-32 Inkling adapter in two passes.” Visual: 50 → 42 → 84 steps.

22–38s

Boundary

“Inkling does not make the voice. It writes performance direction. Fish will render both models with the same voice.” Visual: Inkling text → Fish audio.

38–56s

Twist

“The adapter exists, but I trained mostly script-to-markup—not product-brief-to-great-ad. So a flashy Red Bull prompt is not the clean test.”

56–74s

Evaluation

“The real test locks one script, preserves every word, renders both outputs blind, and reveals identity only after I choose.”

74–90s

Close

“Pipeline proven. Improvement pending. That is the result—and the next experiment is better because of it.”

04 / Safety

Screen-recording safety and publishing checklist

Use the public site, synthetic examples, green test summaries, abstract model diagrams, and redacted dashboards. Do not improvise a terminal tour while authenticated provider state is visible.

Safe to showKeep off-screen
Aggregate counts, task structure, sanitized UI, test resultsEnvironment files, API values, request headers
Generic base → adapter → renderer diagramsPrivate checkpoint, session, or sampler identifiers
Synthetic scripts and expressive examplesRaw ads, client names, source URLs, private transcripts
“Checkpoint saved” and step totalsLocal machine paths, hashes, provider logs
Pending comparison interfaceA fabricated or prematurely revealed A/B winner
Title and thumbnail directions

Title: “I Trained an Audio Model on 42 UGC Ads—Did It Learn the Delivery?”

Alternative: “Can AI Hear What Makes a UGC Ad Work?”

Thumbnail: one waveform on the left, one expressive script on the right, and the restrained line “PIPELINE ≠ PROOF.” Avoid model logos and exaggerated result claims.

Publishable disclosure

This is an independent experiment using a private ad library and hosted training services. Source media, credentials, and checkpoint identities are not public. Training completion is verified; expressive-quality improvement remains pending a blind listening evaluation.