AI 声音经济全景图 · 一次播放的分账 · 2026-07 · 音乐×配音×播客×现场The Voice Economy × AI · The Price of One Play · Jul 2026 · Music × Dubbing × Podcasts × Live

声音被清扫到哪一层? How deep does the sweep go?

AI 对声音产业的冲击不是「要不要来」,而是「清扫到哪一层」:它从最标准化、最没有人格的一端开始吞噬——功能性音乐、译制配音、有声书朗读先倒下;越靠近「肉身在场」与「人格崇拜」的一端,护城河越深。整个产业的价值重心,正从「录音室里可复制的比特」迁往「现场不可复制的在场 AI’s assault on sound is not a question of «whether» but of «how deep the sweep goes»: it devours from the most standardised, least personal end — functional music, dubbing and audiobook narration fall first — while the moat deepens toward bodily presence and the cult of the person. The industry’s centre of value is migrating from the studio’s copyable bits to the stage’s uncopyable presence

本图相信:两根轴可以定位一切——功能性↔人格性(可替换的背景音 vs 你爱的这个人)与录音室↔现场(可无限复制的录音 vs 一次性的肉身在场)。AI 的清扫顺序严格沿轴推进:功能性先死,人格性后死;录音室贬值,现场升值。产业里每一个正在发生的变化,都能在这两根轴上找到坐标。This map believes: two axes locate everything — functional ↔ personal (replaceable background sound vs the person you love) and studio ↔ stage (infinitely copyable recordings vs one-off bodily presence). AI sweeps strictly along them: the functional dies first, the personal last; the studio devalues as the stage appreciates. Every change under way in this industry has coordinates on these two axes.
本图不相信:「听不出=接受」。97% 听众已无法分辨 AI 与人类音乐、专业人士也分不清克隆语音——但一旦被告知是 AI,接受度骤降(消费者净兴趣半年内从 −13% 跌至 −20%,年轻世代降幅最大;62% 表示不愿参与 AI 音乐)。技术门槛已破,关系门槛未破:AI 攻不下的不是「像不像」,是「信不信、爱不爱、在不在场」。It does not believe «indistinguishable means accepted». 97% of listeners can no longer tell AI music from human, and professionals fail on cloned voices — yet told it is AI, acceptance craters (net consumer interest fell from −13% to −20% in six months, steepest among the young; 62% want no part of AI music). The technical threshold has fallen; the relational one has not: what AI cannot take is not «does it sound alike» but «do I trust, do I love, is anyone there».

最小单元:一次播放的分账——流媒体把音乐定价到千分之几美元一次,这道数学题是全篇的诚实层。主脊:一段声音的生命周期七环节(创作制作→登记分发→播放分账→克隆与授权→配音平行轨→播客电台→现场与长尾),③④⑤标红;三条结构带:分账数学、诉讼到收租、零号病人谱系(翻译→原画→电商模特→配音);判断层:人格与现场 > 录音室作品;老目录 > 新作赌博。姊妹:影视分身→film,演出→show,翻译(谱系第一环)→translate,传媒→media The atom: the payout of one play — streaming priced music at thousandths of a dollar a spin, and that arithmetic is the whole map’s honesty layer. The spine: one sound’s lifecycle in seven links (creation → registration and distribution → the play payout → cloning and licensing → the dubbing parallel track → podcasts and radio → the stage and the long tail), ③④⑤ flagged; three bands: the payout maths, from lawsuits to rent, the patient-zero lineage (translation → illustration → e-commerce models → dubbing); the judgment: person and stage > studio work; old catalogues > new-release gambles. Siblings: film replicas → film, live shows → show, translation (the lineage’s first link) → translate, media → media.

$0.003–0.005
一次播放的分账基准(最常被引用的平台,事后估算区间;各平台从 ~$0.001 到 ~$0.015 不等)——千分之几美元的原子定价,是理解整个声音经济的起点(C 估算)The benchmark payout of one play (the most-cited platform, post-hoc estimates; platforms span ~$0.001 to ~$0.015) — atomic pricing in thousandths of a dollar is where understanding this economy begins (C, estimates)
21%
美国配音员称「直接因 AI 失去工作」的比例(1379 份样本,一年内从 14% 升至 21%);中国短剧真人配音订单仅剩约 10%;全球逾 200 万配音从业者面临风险(B)——配音是订单坍塌的「进行时」US voice actors reporting work «lost directly to AI» (1,379 respondents; 14% → 21% in a year); real-voice orders in Chinese micro-drama dubbing down to ~10%; 2M+ practitioners at risk worldwide (B) — dubbing’s collapse is present-tense
44%
Deezer 2026-04 每日上传中纯 AI 生成曲目的占比(约 7.5 万首/天;2025 初仅 ~10%)——而其中高达 85% 的播放为刷量欺诈,合规消费仅占总播放 1–3%(A/B)Share of purely AI-generated tracks in Deezer’s daily uploads by Apr 2026 (~75k a day; ~10% in early 2025) — of which 85% of plays are streaming fraud, legitimate listening at 1–3% of the total (A/B)
90 亿美元
唱片业对 Suno 的理论最高索赔(每部作品最高 15 万美元 × 追加后 6.1 万首)——随后剧情反转:两大厂牌先后和解、转为「授权收租」,老目录成为流媒体+AI 训练的双重收租资产(A/B)The record industry’s theoretical maximum claim against one AI music firm ($150k per work × 61k works after the amendment) — then the twist: two majors settled in turn, converting litigation into licensing rent, and the old catalogue into a double-rent asset (streaming + AI training) (A/B)
⚠️ 口径裁判(先读):① 单次播放费率均为事后估算,平台不公开固定单价,不同来源区间差异大,只作数量级参考;② 全球录制音乐盘子两口径不可直比(一行业组织口径 2025 年 317 亿美元 vs 一研究机构更宽口径 2024 年 362 亿——年份与方法学均不同);③ 「AI 音乐市场规模」各机构从 5.7 亿到 60+ 亿美元不等,本页不采信单一数字;④ 中国配音行业数据多来自媒体报道(订单剩 10%、收入降八成),缺权威普查、可能个案放大;⑤ 「AI 报酬 85 倍」「13% 幻觉率」等二手数字存疑未采;⑥ 诉讼、和解与估值时效性强(截至 2026-07 仍在变动);⑦ 四份深度研究交叉整理(跨代理一致性已核:分账、诉讼、AI 占比、估值轨迹);渗透%为编辑估值。 ⚠️ Basis rulings (read first): ① per-play rates are post-hoc estimates — platforms publish no fixed price, sources vary widely; magnitude only; ② the global recorded-music pie carries two incomparable bases (IFPI’s $31.7B for 2025 vs MIDiA’s wider $36.2B for 2024 — different years and methods); ③ «AI music market» sizings run $570M to $6B+ across houses — no single figure adopted; ④ China dubbing data comes mostly from press reports (orders at 10%, income down 80%) without an authoritative census — case amplification possible; ⑤ second-hand figures («85× AI pay», «13% hallucination rate») are doubted and unused; ⑥ suits, settlements and valuations move fast (current to Jul 2026); ⑦ cross-compiled from four deep-research reports (cross-agent consistency verified on payouts, litigation, AI shares, valuations); penetration %s are editorial.
中心装置 · 两根轴The central device · the two axes
功能性先死,人格性后死;录音室贬值,现场升值The functional dies first, the personal last; the studio devalues, the stage appreciates
判断 AI 渗透到哪里,两根轴即可定位:一段声音是「可替换的功能」还是「不可替换的人格」;价值在「可无限复制的录音」里还是在「一次性的肉身在场」里。AI 能量产录音室比特,却无法复制在场。Two axes locate any incursion: is a sound a replaceable function or an irreplaceable person; does its value live in infinitely copyable recordings or in one-off bodily presence. AI can mass-produce studio bits; it cannot copy presence.
已被清扫的(功能端)Swept (the functional end)
功能性音乐先倒:Spotify 被曝与制作方合作生产低成本「合乎口味」内容(PFC)投放助眠/专注/白噪音歌单以压版税(瑞典调查:91 名制作人产出 1.3 万首);纯 AI 曲目占每日上传 44%。译制/有声书/短剧配音坍塌进行时:AI 朗读作品破 4 万部、生成式语音低至 $30/百万字符、真人订单剩一成(A/B)。 Functional music fell first: Spotify was exposed commissioning cheap «fit-for-taste» content into sleep, focus and white-noise playlists to depress royalties (a Swedish investigation: 91 producers behind 13k tracks); pure-AI tracks reach 44% of daily uploads. Dubbing and audiobooks collapse in real time: 40k+ AI-narrated titles, generative speech from $30 per million characters, live orders down to a tenth (A/B).
啃不动的(人格端)Unchewed (the personal end)
现场收入历史新高(一巡演巨头年收 252 亿美元、1.59 亿人次;一场巡演总票房 22 亿美元);主播口播广告 CPM 20–80 美元、是程序化的数倍——差价买的是「信任转移」;全息演唱会 1.38 亿美元年收的前提是本尊经 5 周动捕全程主导,家属授权的已故歌手全息巡演则遭乐迷批「廉价剥削」(A/B)。 Live revenue sets records (the touring giant at $25.2B and 159M attendees; one tour grossing $2.2B); host-read ads price at $20–80 CPM, multiples of programmatic — the spread buys transferred trust; the hologram residency’s $138M year rested on the artists themselves leading five weeks of motion capture, while an estate-licensed hologram tour of a late singer drew «cheap exploitation» (A/B).
判词Verdict技术门槛已破,关系门槛未破。97% 听不出 AI,但被告知即拒(净兴趣 −20%)——AI 攻不下的不是「像不像」,是「信不信、爱不爱、在不在场」。未来 3–5 年,产业价值将加速向三个不可合成的锚点集中:人格、现场、信任——恰是终章三大母题在声音业的坐标。The technical threshold has fallen; the relational one stands. 97% cannot hear the difference, yet disclosure repels (net interest −20%) — what AI cannot take is not likeness but trust, love and presence. Over the next three to five years value will concentrate on three unsynthesisable anchors: the person, the stage, the trust — precisely the finale’s motifs, plotted in sound.
Reading the Map

从这张图带走的五条规律Five patterns to take away

立场声明:本页是批判性、祛魅的行业结构分析,用 A–D 角标区分官方文件/判决/财报、带源研究、事后估算与厂商口径;费率与市场规模只作数量级;两套全球盘子口径并列;存疑二手数字未采。本页提供行业结构信息,不构成职业、投资或法律建议。核心判断一句话:AI 沿「功能性→人格性、录音室→现场」两根轴逐层清扫——分账数学被压垮的是腰尾部,现场与信任在升值;嗓音正在资产化,头部收租、腰尾被合成;技术门槛已破,关系门槛未破。 Stance: a critical, demystifying structural analysis; A–D badges separate official documents/judgments/filings, sourced research, post-hoc estimates and vendor claims; rates and market sizes are magnitudes; both global bases shown; doubtful second-hand figures unused. Structure only — no career, investment or legal advice. One line: AI sweeps layer by layer along two axes — functional to personal, studio to stage; the payout maths crushes the middle and tail while the stage and trust appreciate; the voice is turning into an asset, the top collecting rent as the rest gets synthesised; the technical threshold has fallen and the relational one has not.