From camera to screen: Forging Vulkan-accelerated codecs in FFmpeg
カメラからスクリーンへ — FFmpegでVulkan高速化コーデックを鍛え上げる
Lynne — FFmpeg Tokyo Video Tech #13
0.5 / about me
About me
自己紹介
Physicist.
物理学者。
FFmpeg developer. Wrote and maintain the Vulkan code, audio encoders and decoders, and the FFT / DCT code.
FFmpeg開発者。Vulkanコード、音声エンコーダー/デコーダー、FFT / DCTコードを執筆・保守。
Member of Khronos, VideoLAN, and the Alliance for Open Media.
Khronos、VideoLAN、Alliance for Open Mediaのメンバー。
Filmmaker.
映像作家。
(thanks, AI)
1 / so you want to make a film
So you want to make a film?
映画を作りたい?
A good script, direction and production help. But before all else, you need a camera.
良い脚本、演出、制作は大事。でも何より先に、カメラが必要。
Adopted by the Library of Congress, among many other archives
米国議会図書館をはじめ、多くのアーカイブ機関が採用
Luckily, I'm a Khronos member.
幸い、私はKhronosのメンバーだ。
I thought about it for a bit:
少し考えてみた:
Me 私
“I mean, I guess I can parallelize Golomb mode using this neat algorithm I stayed up all night designing…”
「まあ、一晩中かけて設計したこの素敵なアルゴリズムでGolombモードなら並列化できると思うけど…」
Them 先方
“No. We want range coding.”
「いや、レンジ符号化がいい。」
Me 私
“That's impossible.”
「それは不可能だ。」
8 / ffv1
After some convincing…
説得の末…
FFv1 allows up to 1024 slices per frame. Each has its own range coder and context state. And the prediction step parallelizes trivially.
FFv1は1フレームあたり最大1024スライスを許容する。各スライスは独自のレンジコーダーとコンテキスト状態を持ち、予測ステップは容易に並列化できる。
9 / the range coder
The range coder cannot be parallelized
レンジコーダーは並列化できない
But 31 helper invocations can do everything around it: preload the context state, apply the state updates, run the RCT, write the pixels.
しかし31個の補助インボケーションがその周りの全てを担える:コンテキスト状態の先読み、状態更新の適用、RCTの実行、画素の書き出し。
9.1 / hybrid decoders
Why hybrid decoders never work
ハイブリッドデコーダーがうまくいかない理由
The round trip is too expensive.
往復のコストが高すぎる。
Coefficients are usually wider than 8 bits. They cost more to upload than the finished decode would.
係数はたいてい8bitより幅が広い。完成した復号結果よりアップロードのコストが高い。
dav1d tried to move as much post-processing as possible onto the GPU. It never beat pure CPU decoding. Not even on low-power devices.
dav1dは後処理をできる限りGPUへ移そうとした。純粋なCPUデコードに勝てなかった。低消費電力デバイスでさえ。
9.5 / ffv1 decode, end to end
FFv1 decoding on the GPU
GPUでのFFv1デコードの全体像
10 / results
Benchmark results
ベンチマーク結果
500 fps
decode — 4K, 10-bit · RTX 6000 Ada デコード — 4K・10bit・RTX 6000 Ada
2.5 → 30 fps
6K×5K 16-bit — Alder Lake monster CPU → cheap 7900 XTX 6K×5K 16bit — Alder Lakeの化け物CPU → 安価な7900 XTX
FFv1, this laptop FFv1、このノートPC
i5-1345U CPU・12スレッド
RX 6900 XT 外付けGPU
Range coder, 1024 slices. 4K clip transcoded from a 750 Mbps ProRes source — noisy content, worst case for a range coder.
レンジコーダー、1024スライス。4Kクリップは750MbpsのProResから変換 — ノイズの多い素材で、レンジコーダーにとって最悪のケース。
Even a codec with no good parallelization method gains enormously.
優れた並列化手法のないコーデックでさえ、得られる利益は莫大。
But decoding must happen on the GPU.
しかし、デコードはGPU上で行わなければならない。
A 5.7K 16-bit RGB frame is 105 MB. At 30 fps: 3 GB/s. Each way.
5.7K 16bit RGBのフレームは105MB。30fpsなら3GB/s。片道で。
PCIe transfers: overhead, latency, CPU load
PCIe転送:オーバーヘッド、レイテンシ、CPU負荷
Not to mention: CPUs stall
言うまでもなく:CPUはストールする
10.4 / results, someone else's machine
Third-party numbers
第三者による計測
FFv1, × realtime
GPU dec
CPU dec
GPU enc
CPU enc
RTX 5070 Ti vs i7-13700K, DPX ⇄ FFv1, FFmpeg 8.1, report dated 2026-08-05. Same slice count on both sides; SSD-bound at the low end.
RTX 5070 Ti対i7-13700K、DPX⇄FFv1、FFmpeg 8.1、2026年8月5日付の報告。高解像度側ではSSD速度が上限。
10.5 / beyond archival
FFv1, beyond archival
アーカイブの先へ
The optimizations let FFv1 move past its archival niche: one format, every stage.
この最適化により、FFv1はアーカイブ用途の枠を超えた:ひとつのフォーマットで全工程を。
11 / back to prores
Now, back to ProRes
さて、ProResに戻ろう
Better than FFv1: the entropy coder is a Golomb-family code. Every lane can guess a codeword. Proper parallel decoding — and encoding.
FFv1より有利:エントロピー符号がGolomb系。全レーンがコードワードを「推測」できる。本格的な並列デコード — そしてエンコードも。
12 / results
Benchmark results
ベンチマーク結果
ProRes HQ ProRes HQ
i5-1345U CPU・12スレッド
Iris Xe 内蔵GPU
RX 6900 XT 外付けGPU
RTX 6000 Ada ワークステーションGPU
13 / back to the camera
So, back to the camera
さて、カメラに戻ろう
You'd like to record in RAW. Cameras don't know which mix of red, green and blue is “white”. And you're too poor for sets and good lighting.
RAWで記録したい。カメラは赤・緑・青のどの組み合わせが「白」なのか知らない。それにセットも良い照明も買えない。
So you get a Panasonic Lumix S1II. Its sensor performs like a $50,000 ARRI's.
そこでPanasonic Lumix S1IIを買う。センサーは5万ドルのARRI並みの性能。
Apple has you covered! At a price — despite already paying for it in the camera. And on a limited operating system.
Appleがカバーしてくれる!有料で — カメラ代に含まれているはずなのに。しかも限られたOSでのみ。
14 / tools
A binaryバイナリ
+
GhidraGhidra
=
🪄
Magic魔法
15 / digging
You dig into the code, tracing where every bit of the input ends up…
コードを掘り下げ、入力の全ビットがどこへ行き着くのかを追跡する…
16 / first light
Then you get lost
そして迷子になる
But you write something good enough to see an image. It doesn't look right.
それでも画像が見える程度のものは書ける。正しくは見えないが。Eventually, something right comes out.
やがて、正しいものが出てくる。
Ship it. We'll figure it out later. Users can monitor with it.
出荷だ。細かいことは後で。ユーザーはモニタリングに使える。
17 / the gpu version
You look around. There's a GPU version.
周りを見渡すと、GPU版がある。
But maybe the GPU version is cleaner…
でもGPU版の方がコードは綺麗かもしれない…
…oh wait.
…あ、待って。
18 / the kernel
The CUDA kernel is encrypted
CUDAカーネルは暗号化されている
XOR with a 64-byte key. Somewhere. Good luck.
64バイトの鍵とのXOR。どこかに。ご武運を。
Claude found it in a minute.
Claudeは1分で見つけた。
19 / the full picture
It's not great
出来は良くない
A CPU decoder bolted onto a GPU DCT. Only for the data to travel back to the CPU.
GPU DCTにCPUデコーダーをくっつけたもの。データはまたCPUに戻っていく。
But thanks to the decrypted kernel, we now have the full picture
しかし復号されたカーネルのおかげで、全体像が見えた
20 / results
ProRes RAW on the GPU
GPU上のProRes RAW
Decode デコード
i5-1345U CPU・12スレッド
Iris Xe 内蔵GPU
RX 6900 XT 外付けGPU
Camera → GPU → screen. No Mac. No accelerator card. No licence.
カメラ → GPU → スクリーン。Macなし。アクセラレータカードなし。ライセンスなし。
Same shaders on Intel, AMD, NVIDIA. Windows, Linux, macOS.
Intel・AMD・NVIDIAで同じシェーダー。Windows・Linux・macOS。
Lives in FFmpeg. Ships everywhere FFmpeg ships.
FFmpeg本体に統合 — FFmpegが届く所すべてに。
21 / trust
Can you trust them?
信頼できるのか?
All Vulkan codecs are confirmed and monitored to match the software implementations.
全てのVulkanコーデックは、ソフトウェア実装と一致することが確認され、継続的に監視されている。
The ProRes implementations are unofficial. They were reverse engineered.
ProRes実装は非公式。リバースエンジニアリングで作られた。
There are no specifications. Even if you speak Chinese.
仕様書は存在しない。中国語が読めたとしても。
Their output is monitored and maintained to match the official implementations.
出力は公式実装と一致するよう監視・維持されている。
I trust them myself. I wrote them. And the binary is so thoroughly reverse engineered that there are no secrets left.
私自身は信頼している。私が書いたからだ。そしてバイナリは徹底的にリバースエンジニアリングされ、秘密はもう残っていない。
(Also, they're literally simpler than JPEG. I can design and write a better codec in one hour. Entirely in GNU nano. With zero AI.)
22 / apv
Enter APV
APVの登場
Obviously, the ProRes situation is quite monopolistic. The IETF agreed. So did Samsung. And Google.
ProResの状況が独占的なのは明らか。IETFも同意した。Samsungも。Googleも。
An IETF-standardized mezzanine codec. Like ProRes, but better.
IETFで標準化されたメザニンコーデック。ProResのようで、ProResより良い。
Measured today on this laptop's eGPU. 1×1 MB tiles are 8,160 tiles per frame, below the spec minimum, and the decoder scales with them.
本日このノートPCのeGPUで計測。1×1 MBタイルは1フレーム8,160タイルで仕様上の最小値未満、デコーダーはタイル数に比例して速くなる。
25 / dpx
DPX
DPX
An SMPTE-standardized raw container. The archival community called back.
SMPTEで標準化された非圧縮コンテナ。アーカイブ業界からまた連絡が来た。
Issue: DPX files produced by film scanners are massive.
問題:フィルムスキャナーが出すDPXファイルは巨大。
51 MB
one 4K film scan frame, 10-bit RGB 4Kフィルムスキャン1フレーム、10bit RGB
1.2 GB/s
to play it at 24 fps 24fpsで再生するのに必要な帯域
8.8 TB
a two-hour film. 13 TB at 16-bit 2時間の映画。16bitなら13TB
26 / dpx, inside
DPX file syntax
DPXファイルの構文
Vendors themselves don't produce standardized files.
ベンダー自身が標準に沿ったファイルを作っていない。
27 / jpeg 2000
JPEG 2000
JPEG 2000
What every cinema worldwide uses.
世界中のあらゆる映画館が使っているもの。
The slowest entropy decoder ever designed. Hostile to CPUs.
史上最も遅いエントロピーデコーダー。CPUに敵対的。
Conceived as an internet-first, progressive codec. The internet never caught on: worse efficiency than JPEG.
インターネット向けのプログレッシブなコーデックとして構想された。インターネットには受け入れられなかった:JPEGより効率が悪い。
Cameras these days produce 48-megapixel JPEGs. 50 MB each. Or more.
最近のカメラは4800万画素のJPEGを出す。1枚50MB。それ以上のことも。
It would be nice to preview them quickly.
サッとプレビューできたら嬉しい。
Hardware JPEG decoders in consumer products are awful.
民生品のハードウェアJPEGデコーダーはひどい。
Often 4:2:0 and 4:2:2 only
4:2:0と4:2:2しか扱えないことが多い
And slow
そして遅い
31 / jpeg, inside
At first glance, JPEG decoding is hopelessly serial
一見、JPEGのデコードは絶望的に直列
But Huffman codes have an interesting property
しかしハフマン符号には面白い性質がある
They resynchronize.
自己同期する。
32 / parallel jpeg
Parallel JPEG decoding
JPEGの並列デコード
Lumix JPEG, 4:2:2 Lumix JPEG・4:2:2
i5-1345U CPU
Iris Xe 内蔵GPU
RX 6900 XT 外付けGPU
Weißenberger & Schmidt, ICPP 2018, modelled on the jpeggpu CUDA reference. The iGPU only guarantees 8-lane subgroups; the decoder needs 32.
Weißenberger & Schmidt(ICPP 2018)、CUDA実装jpeggpuを参考に。内蔵GPUは8レーンのサブグループしか保証せず、デコーダーは32レーンを要求。
33 / state
Where things are
現状
JPEG 2000 decoder: seeking funding to merge
JPEG 2000デコーダー:マージのための資金を募集中
JPEG decoder and encoder: seeking funding to merge
JPEGデコーダー・エンコーダー:マージのための資金を募集中
DNxHD decoder and encoder: seeking funding to finish
DNxHDデコーダー・エンコーダー:完成のための資金を募集中
Contact: dev@lynne.ee連絡先:dev@lynne.ee
(filmmaking isn't cheap)
34 / the point
The point of all this
この仕事の目的
Everyone gets to work with high-quality video on consumer GPUs. On any platform. Freely.
誰もが民生GPUで高品質な映像を扱える。どのプラットフォームでも。自由に。
Every Vulkan codec uses FFmpeg's native hardware acceleration framework
全てのVulkanコーデックはFFmpeg標準のハードウェアアクセラレーション基盤を使用
Instant switching between GPU and CPU decoding. API users enable it with one lineGPUデコードとCPUデコードを即座に切り替え可能。APIの利用者は1行で有効化できる
mpv全コーデックを自動で対応済み
mpv already supports every one of them, automatically
mpvは既に全コーデックに自動対応
Shotcut, the NLE, uses FFmpeg hardware decoding and encoding
NLEのShotcutはFFmpegのハードウェアデコード・エンコードを利用
Blender: a pull request adds the Vulkan codecs to the video editor
Blender:Vulkanコーデックを動画編集機能に加えるプルリクエストあり