INKLINGUGC PERFORMANCE LAB

PUBLIC BUILD / 2026

03 / Dataset

50 selected / 42 accepted

From 3,000 ads to a reviewable cohort

The best ads in a library are not automatically the best training rows. Selection, transcription, expressive labeling, and quarantine each protect a different part of the signal.

Selection, transcription, expressive labels, quarantine, and the final training mix.

01 / Selection

A human chose the performances first

The project began with more than 3,000 locally stored UGC ads collected as strong creative references. A local review dashboard streamed one ad at a time and offered a simple decision: train or do not train. Fifty performances were frozen as the first cohort.

Library3,000+Private source ads
Selected50Frozen cohort
Accepted42Training rows
Excluded8Quarantined

Performance metrics were intentionally absent from the review interface. The decision was about audible creative value, not a hidden proxy such as spend or click-through rate.

02 / Labels

Every target joined words to delivery

Each accepted example bound a source transcript to expressive direction. The useful parts were global tone, pacing, energy arc, segment-level timing, and a performance-ready version of the same words.

Global

Direction

The overall emotional posture, pacing, and energy arc across the ad.

Local

Segments

Source spans paired with timing and delivery notes at meaningful boundaries.

Constraint

Same words

Expressiveness could change the performance, not silently rewrite the copy.

Why wording preservation mattered

If the target changes both copy and performance, a later difference is hard to interpret. The model may have learned rewriting, markup, or both. Holding spoken words fixed creates a clearer supervised relationship.

03 / Training mix

Most rows were text-only by design

The final 42 accepted rows included 31 text-only examples and 11 audio-conditioned examples. That mix reflected the real training objective: produce useful expressive markup when only a script is available, while retaining some direct audio grounding.

31 rows / 74%

Text-only

Input provides the source script and task context. The target contains the same words with expressive direction.

11 rows / 26%

Audio-conditioned

Input also includes source audio, allowing the label to be grounded directly in heard delivery.

Interpretation

The audio-conditioned rows help connect sound to language. The text-only majority aligns the adapter with the intended inference-time task.

04 / Quarantine

Eight rows stayed out instead of being hand-waved in

Rows were quarantined when the preparation or target contract could not be trusted. Typical causes included incomplete or malformed output, a mismatch between source words and expressive text, problematic media, or a label that could not pass the final schema checks.

Media and transcript readiness

Confirm the chosen record can be processed and the transcript represents its spoken content.

Teacher output integrity

Require parseable, structurally complete expressive direction rather than repairing arbitrary prose silently.

Spoken-word preservation

Detect additions, deletions, or substitutions in words that should have stayed fixed.

Training render

Reject rows that cannot become a deterministic supervised example.

Quarantine is not wasted data. It is a record of where the labeling process needs to improve before those examples are reconsidered.