Generate Ads
Enter a product, audience, verified facts, offer, duration, and creative direction. Both models write a complete ad. Differences may come from copywriting, structure, or expressive markup.
PUBLIC BUILD / 2026
The public question is simple—did the adapter make expressive direction better? The answer requires matched prompts, preserved words, and ears, not just two impressive-looking text boxes.
Two comparison modes and the listening test that still needs to happen.
01 / Two modes
A local comparison dashboard can load Base Inkling and the final trained adapter under the same generation settings. It supports two modes because the project is interested in both creative usefulness and the narrower behavior represented in training.
Enter a product, audience, verified facts, offer, duration, and creative direction. Both models write a complete ad. Differences may come from copywriting, structure, or expressive markup.
Paste one approved script. Both models may add performance direction, but every spoken word must remain unchanged. This is the closer test of the supervised task.
Model identity is the intended difference. Prompt, temperature, token limit, and renderer settings remain paired so a comparison is interpretable.
02 / Text stage
Each arm records the raw model response and the extracted performance-ready text. That separation catches a common failure: a model may write commentary around the requested script or return a field name different from the expected one.
| Check | Why it matters | Failure state |
|---|---|---|
| Response extracted | The dashboard found one usable script field without guessing. | Preserve raw response and stop that arm. |
| Spoken words preserved | Controlled mode changes direction, not claims or copy. | Mark the arm invalid for the diagnostic. |
| Expressive cues visible | Reviewers can inspect density, specificity, and placement. | Record an unmarked or malformed response honestly. |
| Both arms complete | Only paired outputs move to the same-voice renderer. | Keep the successful arm; do not fabricate a pair. |
Adding more tags is not automatically better. Useful direction should clarify an audible or intended performance without turning every phrase into competing stage instructions.
03 / Audio stage
Fish Audio S2-Pro is the planned downstream renderer. The same voice reference and synthesis settings will render both texts. The interface randomizes their identity as A and B, withholds the mapping, and asks for A, B, tie, or neither.
Freeze the prompt, scripts, configurations, and model identities.
Only the expressive text changes between audio arms.
Listen before knowing which model produced either script.
Lock the decision, reveal identity once, and freeze the record.
Fish Audio has not yet been called for this comparison. No public page presents a synthetic listening winner.
04 / Verdict
The available text outputs can suggest that the adapter uses different or richer performance language. A small training set can also produce superficial changes, unstable cue placement, or behavior that already existed in the base model.
Base and trained identities can be loaded separately; the trained adapter completed the intended supervised run.
There is not yet a completed, sufficiently sized blind listening evaluation showing the trained arm wins.
The next claim should follow the evidence: either the trained output wins on untouched scripts, it ties the base model, or it loses. Every outcome teaches us something useful about the data design.