From camera to screen:
Forging Vulkan-accelerated codecs in FFmpeg

カメラからスクリーンへ — FFmpegでVulkan高速化コーデックを鍛え上げる
Lynne — FFmpeg Tokyo Video Tech #13
0.5 / about me

About me

自己紹介

Filmmaker. 映像作家。

(thanks, AI)

1 / so you want to make a film

So you want to make a film?

映画を作りたい?

A good script, direction and production help. But before all else, you need a camera. 良い脚本、演出、制作は大事。でも何より先に、カメラが必要。

Cameras create multimedia files. カメラはマルチメディアファイルを生み出す。

Tweet: multimedia is basically neverending pain
multimedia is basically neverending pain マルチメディアとは、基本的に終わりのない苦痛である
2 / high quality video

You need to know how to
work with high-quality video

高品質な映像の扱い方を知っておく必要がある
Generation loss. On a loop. 世代劣化。延々と。
3 / choosing a codec

Choosing a codec

コーデックを選ぶ

In modern times 現代なら

  • You'd likely choose intra-only H.264, or intra-only AV1, CQP=1 おそらくイントラオンリーH.264か、イントラオンリーAV1、CQP=1を選ぶだろう
  • Hardware encoders and decoders everywhere ハードウェアエンコーダー/デコーダーがどこにでもある
  • Hundreds to thousands of fps at 8K 8Kで数百〜数千fps
  • On most laptops. And mobile phones. 大抵のノートPCで。スマートフォンでも。

The standard solution: ProRes 標準的な解決策:ProRes

  • Interoperable with everyone 誰とでも相互運用できる
  • Mastering マスタリング
  • Archival アーカイブ
  • Editing 編集
  • Camera footage カメラ撮影素材

ProRes. Everywhere.

ProRes。あらゆる工程に。

The one step that isn't ProRes takes ProRes as its input. ProResではない唯一の工程も、入力はProRes。

…except if you're on Linux. …Linuxを使っていなければ、の話だが。

ProRes — or rather, the entire professional video ecosystem — locks you into a Mac. ProRes、いや、プロ向け映像エコシステム全体が、Macに囲い込んでくる。

3.5 / the price of admission
Apple Afterburner card, list price US$2,000

The price of admission

入場料
  • A $2,000 accelerator card 2,000ドルのアクセラレータカード
  • An FPGA. PCIe 3.0 x16. Mac Pro only. FPGA。PCIe 3.0 x16。Mac Pro専用。
  • A licensed encoder and decoder from a big vendor 大手ベンダーのライセンス付きエンコーダーとデコーダー
  • …and so on. …などなど。
4 / prores, inside

ProRes is a highly complex and efficient codec

ProResは高度に複雑で効率的なコーデック

A DCT step, a prediction step, a quantization step, an entropy coding step. DCT、予測、量子化、エントロピー符号化の各ステップ。

Wait a minute… ちょっと待って…

It's literally just JPEG. Same 8×8 transform. Simpler prediction. Simpler coding. 文字通りただのJPEG。同じ8×8変換。より単純な予測。より単純な符号化。

5 / better than jpeg

Better than JPEG: embarrassingly parallel

JPEGより優れた点:恥ずかしいほど並列化できる

Thousands of independent slices per 8K frame. And a few CPU cores. 8Kフレームあたり数千の独立したスライス。そしてCPUコアは数個。

And a GPU with literally tens of thousands of lanes, sharing memory and even registers. そしてGPUには文字通り数万のレーンがあり、メモリはおろかレジスタまで共有する。

How to access the GPU? GPUにはどうやってアクセスする?

Vulkan
6.5 / vulkan

Vulkan is a GPU programming API

VulkanはGPUプログラミングAPI
7 / how it started

How FFmpeg started using Vulkan for codecs

FFmpegがコーデックにVulkanを使い始めた経緯
The archival community, 2024 アーカイブ業界、2024年
“FFv1 is too slow to encode. Write CUDA or whatever, just save us!” 「FFv1はエンコードが遅すぎる。CUDAでも何でもいいから、助けてくれ!」

Luckily, I'm a Khronos member. 幸い、私はKhronosのメンバーだ。

I thought about it for a bit: 少し考えてみた:

Me
“I mean, I guess I can parallelize Golomb mode using this neat algorithm I stayed up all night designing…” 「まあ、一晩中かけて設計したこの素敵なアルゴリズムでGolombモードなら並列化できると思うけど…」
Them 先方
“No. We want range coding.” 「いや、レンジ符号化がいい。」
Me
“That's impossible.” 「それは不可能だ。」
8 / ffv1

After some convincing…

説得の末…

FFv1 allows up to 1024 slices per frame. Each has its own range coder and context state. And the prediction step parallelizes trivially. FFv1は1フレームあたり最大1024スライスを許容する。各スライスは独自のレンジコーダーとコンテキスト状態を持ち、予測ステップは容易に並列化できる。

9 / the range coder

The range coder cannot be parallelized

レンジコーダーは並列化できない

But 31 helper invocations can do everything around it: preload the context state, apply the state updates, run the RCT, write the pixels. しかし31個の補助インボケーションがその周りの全てを担える:コンテキスト状態の先読み、状態更新の適用、RCTの実行、画素の書き出し。

9.1 / hybrid decoders

Why hybrid decoders never work

ハイブリッドデコーダーがうまくいかない理由

The round trip is too expensive. 往復のコストが高すぎる。

9.5 / ffv1 decode, end to end

FFv1 decoding on the GPU

GPUでのFFv1デコードの全体像
10 / results

Benchmark results

ベンチマーク結果
500 fps
decode — 4K, 10-bit · RTX 6000 Ada
デコード — 4K・10bit・RTX 6000 Ada
2.5 → 30 fps
6K×5K 16-bit — Alder Lake monster CPU → cheap 7900 XTX
6K×5K 16bit — Alder Lakeの化け物CPU → 安価な7900 XTX
FFv1, this laptop FFv1、このノートPCi5-1345U
CPU・12スレッド
RX 6900 XT
外付けGPU

Range coder, 1024 slices. 4K clip transcoded from a 750 Mbps ProRes source — noisy content, worst case for a range coder. レンジコーダー、1024スライス。4Kクリップは750MbpsのProResから変換 — ノイズの多い素材で、レンジコーダーにとって最悪のケース。

Even a codec with no good parallelization method gains enormously. 優れた並列化手法のないコーデックでさえ、得られる利益は莫大。

But decoding must happen on the GPU. しかし、デコードはGPU上で行わなければならない

10.4 / results, someone else's machine

Third-party numbers

第三者による計測
FFv1, × realtimeGPU decCPU decGPU encCPU enc

RTX 5070 Ti vs i7-13700K, DPX ⇄ FFv1, FFmpeg 8.1, report dated 2026-08-05. Same slice count on both sides; SSD-bound at the low end. RTX 5070 Ti対i7-13700K、DPX⇄FFv1、FFmpeg 8.1、2026年8月5日付の報告。高解像度側ではSSD速度が上限。

10.5 / beyond archival

FFv1, beyond archival

アーカイブの先へ

The optimizations let FFv1 move past its archival niche: one format, every stage. この最適化により、FFv1はアーカイブ用途の枠を超えた:ひとつのフォーマットで全工程を。

FFv1 Camera RAW Editing Mezzanine Mastering Archival Streaming VFX / interchange カメラRAW 編集 メザニン マスタリング アーカイブ 配信 VFX・素材交換
11 / back to prores

Now, back to ProRes

さて、ProResに戻ろう

Better than FFv1: the entropy coder is a Golomb-family code. Every lane can guess a codeword. Proper parallel decoding — and encoding. FFv1より有利:エントロピー符号がGolomb系。全レーンがコードワードを「推測」できる。本格的な並列デコード — そしてエンコードも。

12 / results

Benchmark results

ベンチマーク結果
ProRes HQ ProRes HQi5-1345U
CPU・12スレッド
Iris Xe
内蔵GPU
RX 6900 XT
外付けGPU
RTX 6000 Ada
ワークステーションGPU

13 / back to the camera

So, back to the camera

さて、カメラに戻ろう

You'd like to record in RAW. Cameras don't know which mix of red, green and blue is “white”. And you're too poor for sets and good lighting. RAWで記録したい。カメラは赤・緑・青のどの組み合わせが「白」なのか知らない。それにセットも良い照明も買えない。

So you get a Panasonic Lumix S1II. Its sensor performs like a $50,000 ARRI's. そこでPanasonic Lumix S1IIを買う。センサーは5万ドルのARRI並みの性能。

Apple has you covered! At a price — despite already paying for it in the camera. And on a limited operating system. Appleがカバーしてくれる!有料で — カメラ代に含まれているはずなのに。しかも限られたOSでのみ。

14 / tools
+
=
15 / digging

You dig into the code, tracing where every bit of the input ends up… コードを掘り下げ、入力の全ビットがどこへ行き着くのかを追跡する…

Ghidra decompiling PRRawCreateOpenCLProcessor in ProResRAW.dll
16 / first light

Then you get lost

そして迷子になる
First ProRes RAW decode attempt, garbled
But you write something good enough to see an image. It doesn't look right. それでも画像が見える程度のものは書ける。正しくは見えないが。
ProRes RAW decode, correct image
Eventually, something right comes out. やがて、正しいものが出てくる。

Ship it. We'll figure it out later. Users can monitor with it. 出荷だ。細かいことは後で。ユーザーはモニタリングに使える。

17 / the gpu version

You look around. There's a GPU version.

周りを見渡すと、GPU版がある。

But maybe the GPU version is cleaner… でもGPU版の方がコードは綺麗かもしれない…

…oh wait. …あ、待って。

18 / the kernel

The CUDA kernel is encrypted

CUDAカーネルは暗号化されている

XOR with a 64-byte key. Somewhere. Good luck. 64バイトの鍵とのXOR。どこかに。ご武運を。

Claude found it in a minute. Claudeは1分で見つけた。

Claude's session locating the kernel deobfuscator and key
19 / the full picture

It's not great

出来は良くない

A CPU decoder bolted onto a GPU DCT. Only for the data to travel back to the CPU. GPU DCTにCPUデコーダーをくっつけたもの。データはまたCPUに戻っていく。

But thanks to the decrypted kernel, we now have the full picture

しかし復号されたカーネルのおかげで、全体像が見えた

      
    20 / results

    ProRes RAW on the GPU

    GPU上のProRes RAW
    Decode デコードi5-1345U
    CPU・12スレッド
    Iris Xe
    内蔵GPU
    RX 6900 XT
    外付けGPU

    21 / trust

    Can you trust them?

    信頼できるのか?

    All Vulkan codecs are confirmed and monitored to match the software implementations. 全てのVulkanコーデックは、ソフトウェア実装と一致することが確認され、継続的に監視されている。

    Apple: unauthorized codec implementations such as FFmpeg

    The ProRes implementations are unofficial. They were reverse engineered. ProRes実装は非公式。リバースエンジニアリングで作られた。

    There are no specifications. Even if you speak Chinese. 仕様書は存在しない。中国語が読めたとしても。

    Their output is monitored and maintained to match the official implementations. 出力は公式実装と一致するよう監視・維持されている。

    I trust them myself. I wrote them. And the binary is so thoroughly reverse engineered that there are no secrets left. 私自身は信頼している。私が書いたからだ。そしてバイナリは徹底的にリバースエンジニアリングされ、秘密はもう残っていない。

    (Also, they're literally simpler than JPEG. I can design and write a better codec in one hour. Entirely in GNU nano. With zero AI.)

    22 / apv

    Enter APV

    APVの登場

    Obviously, the ProRes situation is quite monopolistic. The IETF agreed. So did Samsung. And Google. ProResの状況が独占的なのは明らか。IETFも同意した。Samsungも。Googleも。

    An IETF-standardized mezzanine codec. Like ProRes, but better. IETFで標準化されたメザニンコーデック。ProResのようで、ProResより良い。

    23 / apv, inside

    APV, inside

    APVの中身

    It is very similar to ProRes. On purpose. ProResと非常によく似ている。意図的に。

    24 / apv results

    APV on the GPU

    GPU上のAPV
    1000 fps
    encode and decode — 1080p 4:2:2 10-bit, smallest tiles, RX 6900 XT
    エンコードもデコードも — 1080p 4:2:2 10bit・最小タイル・RX 6900 XT
    1 lane / tile
    the decoder is one invocation per tile and component: smaller tiles, more lanes
    デコーダーはタイル×成分ごとに1インボケーション:タイルが小さいほどレーンが増える
    1080p 4:2:2 10-bit, RX 6900 XT 1080p 4:2:2 10bit・RX 6900 XT1×1 MB tiles
    タイル 1×1 MB
    16×8 MB tiles (spec minimum)
    タイル 16×8 MB(仕様上の最小)

    Measured today on this laptop's eGPU. 1×1 MB tiles are 8,160 tiles per frame, below the spec minimum, and the decoder scales with them. 本日このノートPCのeGPUで計測。1×1 MBタイルは1フレーム8,160タイルで仕様上の最小値未満、デコーダーはタイル数に比例して速くなる。

    25 / dpx

    DPX

    DPX

    An SMPTE-standardized raw container. The archival community called back. SMPTEで標準化された非圧縮コンテナ。アーカイブ業界からまた連絡が来た。

    Issue: DPX files produced by film scanners are massive. 問題:フィルムスキャナーが出すDPXファイルは巨大

    51 MB
    one 4K film scan frame, 10-bit RGB
    4Kフィルムスキャン1フレーム、10bit RGB
    1.2 GB/s
    to play it at 24 fps
    24fpsで再生するのに必要な帯域
    8.8 TB
    a two-hour film. 13 TB at 16-bit
    2時間の映画。16bitなら13TB
    26 / dpx, inside

    DPX file syntax

    DPXファイルの構文

    Vendors themselves don't produce standardized files. ベンダー自身が標準に沿ったファイルを作っていない。

    27 / jpeg 2000

    JPEG 2000

    JPEG 2000

    What every cinema worldwide uses. 世界中のあらゆる映画館が使っているもの。

    The slowest entropy decoder ever designed. Hostile to CPUs. 史上最も遅いエントロピーデコーダー。CPUに敵対的。

    Conceived as an internet-first, progressive codec. The internet never caught on: worse efficiency than JPEG. インターネット向けのプログレッシブなコーデックとして構想された。インターネットには受け入れられなかった:JPEGより効率が悪い。

    Same image at about 9 KB: original, JPEG 2000, JPEG XR, JPEG, HEIF
    Same picture, ~9 KB each. JPEG 2000, JPEG XR, JPEG, HEIF. 同じ写真、各約9KB。JPEG 2000、JPEG XR、JPEG、HEIF。
    28 / jpeg 2000, inside

    JPEG 2000, inside

    JPEG 2000の中身
    29 / the mq decoder

    The parallel MQ decoder

    並列MQデコーダー

    The key to fast JPEG 2000 decoding. Serial inside a codeblock. Parallel across 26,000 of them. 高速なJPEG 2000デコードの鍵。コードブロック内は直列。26,000個のコードブロック間で並列。

    181 fps
    2K DCP, RX 6900 XT, this laptop
    2K DCP、RX 6900 XT、このノートPC
    28 fps
    same clip, i5-1345U, 12 threads
    同じクリップ、i5-1345U
    113 fps
    4K DCP, desktop AMD GPU (2K: 215 fps)
    4K DCP、デスクトップAMD GPU
    30 / jpeg

    The original JPEG

    元祖JPEG

    Cameras these days produce 48-megapixel JPEGs. 50 MB each. Or more. 最近のカメラは4800万画素のJPEGを出す。1枚50MB。それ以上のことも。

    It would be nice to preview them quickly. サッとプレビューできたら嬉しい。

    Hardware JPEG decoders in consumer products are awful. 民生品のハードウェアJPEGデコーダーはひどい。

    31 / jpeg, inside

    At first glance, JPEG decoding is hopelessly serial

    一見、JPEGのデコードは絶望的に直列

    But Huffman codes have an interesting property

    しかしハフマン符号には面白い性質がある

    They resynchronize. 自己同期する。

    32 / parallel jpeg

    Parallel JPEG decoding

    JPEGの並列デコード
    Lumix JPEG, 4:2:2 Lumix JPEG・4:2:2i5-1345U
    CPU
    Iris Xe
    内蔵GPU
    RX 6900 XT
    外付けGPU

    Weißenberger & Schmidt, ICPP 2018, modelled on the jpeggpu CUDA reference. The iGPU only guarantees 8-lane subgroups; the decoder needs 32. Weißenberger & Schmidt(ICPP 2018)、CUDA実装jpeggpuを参考に。内蔵GPUは8レーンのサブグループしか保証せず、デコーダーは32レーンを要求。

    33 / state

    Where things are

    現状
    34 / the point

    The point of all this

    この仕事の目的

    Everyone gets to work with high-quality video on consumer GPUs. On any platform. Freely. 誰もが民生GPUで高品質な映像を扱える。どのプラットフォームでも。自由に。

    • mpv already supports every one of them, automatically mpvは既に全コーデックに自動対応
    • Shotcut, the NLE, uses FFmpeg hardware decoding and encoding NLEのShotcutはFFmpegのハードウェアデコード・エンコードを利用
    • Blender: a pull request adds the Vulkan codecs to the video editor Blender:Vulkanコーデックを動画編集機能に加えるプルリクエストあり

    Discussion

    質疑応答
    Lynne — FFmpeg
    1 /
    ← → · Space · F11 fullscreen · Ctrl+P → PDF