Figure 1: Overview of the four-stage SwanData-Caption data processing pipeline, including coverage design, SwanData-Speech preprocessing, caption annotation, and data refinement.
Table 1: Output-field definitions used in the caption annotation schema.
Field
Description
Environment
Scene-level environment and recording context, including location or place description, sound-field impression, room or recording-space cues, microphone characteristics, reverberation, background music, crowd murmur, traffic, wind, rain, electrical hum, keyboard tapping, appliance noise, distant footsteps, or other persistent background sound or effects that function as the scene bed.
Speakers
The inventory of actually speaking subjects. Each speaker is described by perceived gender, age range, persona or role when audible from delivery, stable timbre, articulation, loudness tendency, speaking rate, accent, habitual style, and stable affective tendency.
Content
A chronological content field and fine-grained local style description. Speech is wrapped by speaker tags such as <S1> and </S1>; local audio effects are wrapped by <Audio> and </Audio>. Fine-grained local style, including changes in emotion, volume, pace, pause, emphasis, hesitation, interruption, code-switching, and nearby effect context, is described around the tagged spans.
Figure 2: Overview of SwanTale. Figure (a) shows the architecture of SwanTale, and Figure (b) shows Unified MoE. In (a), the zero-shot path supplies reference audio, while both tasks share text and caption. In (b), a task router selects experts at the sample level, while an audio router applies Top-P routing over frame-level audio and null experts.
Table 3: Reconstruction quality on the speech and singing voice test sets. Bold and underlined values indicate the best and second-best results among the compared systems within each domain, respectively.
Model
PESQ ↑
STOI ↑
MCD ↓
ViSQOL ↑
Speech
DAC [60]
4.1178
0.9693
1.1963
4.1585
EnCodec [19]
3.1872
0.9297
1.5147
3.5035
WavTokenizer Large Unify [48]
2.1423
0.8428
2.9393
2.3787
VoxCPM2 AudioVAE V2 [118]
3.9987
0.9690
1.2222
4.0340
MegaTTS 3 WaveVAE [51]
3.5968
0.9507
1.5130
4.2348
SwanVAE (Ours)
4.1683
0.9680
0.9638
4.1248
Singing Voice
DAC [60]
3.7872
0.8627
1.9293
3.6681
EnCodec [19]
2.6464
0.8166
2.2392
3.3841
WavTokenizer Large Unify [48]
1.9226
0.6613
5.1885
1.6018
VoxCPM2 AudioVAE V2 [118]
3.7088
0.8838
1.9335
3.6734
MegaTTS 3 WaveVAE [51]
3.5727
0.8620
2.0310
4.0013
SwanVAE (Ours)
3.9821
0.9001
1.5661
3.7085
Figure 3: Overview of SwanVAE. (a) The anti-aliased convolutional encoder, Gaussian variational bottleneck, and local Transformer decoder. (b) Generative alignment through flow matching and causal latent prediction. (c) Energy, multi-scale chroma, and multi-band energy readouts with waveform-derived targets. The decoder receives posterior samples 𝐳, while the alignment objectives operate on the posterior mean 𝝁ϕ exclusively during the SwanVAE training stage and do not enter the downstream generator at inference time.
Table 4: Reconstruction quality on the general audio and music test sets. Bold and underlined values indicate the best and second-best results among the compared systems within each domain, respectively.
General Audio
Model
ViSQOL ↑
LSD ↓
DAC [60]
4.0198
0.9589
EnCodec [19]
4.1140
0.9761
WavTokenizer Large Unify [48]
2.8595
1.0967
Stable Audio Open 1.0 [23]
4.0355
0.9358
SAME-L [76]
3.7541
1.0372
SwanVAE (Ours)
4.1269
0.9455
Table 6: Instruct TTS results on InstructTTSEval. Results for all models other than SwanTale are taken from the VoxCPM2 paper [118]. Bold and underlined values indicate the best and second-best results.
Model
Chinese (ZH)
English (EN)
APS ↑
DSD ↑
RP ↑
APS ↑
DSD ↑
RP ↑
Parler-TTS-large [69]
–
–
–
60.0
45.9
31.2
VoxInstruct [117]
47.5
52.3
42.6
54.9
57.0
39.3
VoiceSculptor [39]
75.7
64.7
61.5
–
–
–
MiMo-Audio-7B-Instruct [106]
75.7
74.3
61.5
80.6
77.6
59.5
Qwen3-TTS-12Hz-1.7B-VD [79]
85.2
81.1
65.1
82.9
82.4
68.4
MOSS-VoiceGenerator [42]
78.0
80.0
74.0
68.2
82.0
68.7
VoxCPM2 [118]
85.2
71.5
60.8
84.2
83.2
71.4
SwanTale (Ours)
86.1
80.1
64.1
84.2
79.2
63.6
Table 8: Results on SwanBench-Caption. All metrics are scored on a 1–5 scale by gemini-3.5-flash; higher is better. 32B CE replaces the default Qwen3.0-Instruct-8B caption encoder with Qwen3.0-Instruct-32B [99].
Setting
Instruction Accuracy ↑
Acoustic Quality ↑
Overall Expressiveness ↑
SwanTale w/o MoE
3.02
4.09
3.56
SwanTale
3.39
4.31
3.82
SwanTale w/ 32B CE
3.70
4.34
3.98
Table 9: Condensed style matrix for animation-style captions.
Aspect
Condensed rule
Typical triggers
Animation, cartoon, anime, dubbing, role-playing voices, and other clips whose delivery follows a character-dubbing convention.
Stable speaker profile
Describe perceived gender, approximate age, an audible vocal archetype when useful, stable timbre, and habitual delivery, in that order. Archetypes such as an energetic lead, a restrained mature speaker, or a comic supporting voice require clear evidence in the vocal performance.
Local delivery
Record exaggerated reactions, abrupt emotional shifts, punch-line timing, shouts, laughter, hesitation, and changes in pace, loudness, or arousal in the chronological Content field.
Acoustic evidence
Ground descriptions in cues such as habitual pitch range, brightness, breathiness, energy, attack strength, pausing, and the degree of restraint or exaggeration.
Representative distinctions
Action-oriented clips favor larger loudness dynamics, faster pace, and stronger bursts; romance favors finer emotional control, breathiness, and pauses; suspense favors restrained, clear delivery; historical or courtly settings favor formal diction and measured expression.
Table 10: Condensed style matrix for short-drama and film/TV-drama-style captions.
Aspect
Condensed rule
Typical triggers
Short drama, micro drama, vertical drama, scripted short video, web drama, film, TV drama, and other dialogue-heavy staged media.
Stable speaker profile
Describe perceived gender, approximate age, a role or social identity supported by spoken dialogue or vocal delivery, stable timbre, and habitual delivery.
Local delivery
Record interruption, conflict, emotional escalation, reversal, pleading, threat, command, hesitation, and relationship-driven changes in pace, loudness, or tone in the chronological Content field.
Acoustic evidence
Use audible properties such as pacing, diction, theatrical coloring, controlled pauses, coldness, ingratiating delivery, and abrupt changes in intensity to describe delivery. Character identity requires supporting dialogue or role evidence.
Describe perceived gender, approximate age, a persona type supported by the speech function, voice-style class, timbre, and habitual product-pitch delivery. Persona labels such as host, product recommender, lecturer, or service worker require evidence from the spoken content or delivery.
Local delivery
Record selling-point emphasis, urgency, price or discount emphasis, calls to action, question hooks, trust-building explanations, conversational softening, and changes in excitement in the chronological Content field.
Acoustic evidence
Ground descriptions in audible properties such as friendliness, authority, energy, technical density, conversational warmth, cadence regularity, pause timing, emotional range, and script-like phrasing. Synthetic-voice judgments require direct audible artifacts.
Representative distinctions
Representative groups include product recommenders and livestream hosts, health or education explainers, finance or business speakers, and service roles; matching styles range from conversational sharing and storytelling to energetic sales and structured explanation.
Table 12: Utterance-level accuracy (%) of SwanVerifier on the held-out labeled split.
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, support acoustic scenes with environments and audio effects, and later reuse the designed voices. Therefore, it is important to support multi-speaker speech and audio generation for both instruct and zero-shot tasks. The ins