Let's make the slides for my talk at Tokyo Video Tech. Synopsis was: "Video codecs are designed to serve a purpose, whether its general video compression for the web, or low-latency high-speed/quality for editing, with custom hardware designed for them to run on. For codecs designed to be general-purpose, such hardware is abundant. For professional video codecs, such hardware is highly-specialized, expensive, limited, and proprietary. Computers nowadays come with high-performance general-purpose GPUs, that do less and less graphics. This session will cover the techniques of how custom serial compression meant for special hardware, was parallized and made to run on GPUs. The results were efficient codec implementations, able to run on any platform, thus liberating users from requiring special hardware. Finally, the talk will cover how this technology enabled the entire cinema creation pipeline, from camera to screen, to be performed on regular GPUs using free and open-source tools. Speaker Profile: Lynne is a codec engineer, member of the Alliance for Open Media and Khronos, working on FFmpeg. She worked on AV1, the Vulkan specifications, wrote and maintains multiple encoders, decoders, and assembly for FFmpeg. Currently, she is moving into cinematography and filmmaking." I'm going for a Matthew Garret-like style. Translate points, and captions into Japanese, less noticeable, like in the example presentation provided ("ffv1-vulkan-imagica(2).html", that you yourself made). Feel free to make slide titles. Slide 0: Title: "From camera to screen: Forging Vulkan-accelerated codecs in FFmpeg" Slide 0.5, part 1: about me: Physicist. FFmpeg developer, I wrote and maintain the Vulkan code, audio encoders and decoders, and FFT/DCT code Slide 0.5, part 2: Khronos member, VideoLan member, Alliance of Open Media member Slide 0.5, part 3: Filmmaker Slide 0.5, part 4: (thanks, AI) <- small words Slide 1, part 1: So you want to make a film? Well, it's good to have a good script, direction and production, but before all else, you need a camera. Slide 1, part 2: Cameras create multimedia files. Slide 1, part 3: (caption: multimedia is basically neverending pain) Slide 2: You need to know how to work with high quality video (degradation loss mp4 playback on a loop). Slide 3, part 1: Now, in modern times, you'd likely choose something like intra-only H264, or intra-only AV1, CQP=1, with abundant hardware encoders and decoders, running at hundreds to thousands of FPS at 8k on most laptops and mobile phones. Slide 3, part 2: Or you may go for the standard solution: ProRes. Interoperable with everyone (bullet points: mastering, archival, editing, camera footage). Slide 3, part 3: Slide 3, part 4: ...except if you're on Linux. Slide 3, part 4: Everything in the ProRes, or rather, the entire professional video ecosystem, locks you into a Mac system. Slide 3.5 (sorry, I have to insert a slide, but I already numbered slides afterwards): You'll need a 2000 dollar accelerator card , a licensed encoder and decoder from a big vendor, etc. Slide 4, part 1: ProRes is a highly complex and efficient codec. It's got a DCT step, then a prediction step, then a quantization step, then an entropy encode step. . Wait a minute... Slide 4, part 2: It's literally just JPEG! Same 8x8 transform! Simpler prediction, simpler encoding. Slice 5, part 1: better than JPEG, we can easily massively parallelize it . We have thousands of blocks at 8k. But we only have a few CPU cores. Slide 5, part 2: And we have a GPU with literally tens of thousands of cores, sharing memory and even registers (show average AMD GPU diagram). Slide 6: How to access the GPU? . Slide 6.5: Vulkan is a GPU programming API. Its lower level than CUDA and OpenCL, but its portable, able to run on all plaforms and GPUs, and each implementation is validated. Slide 7, part 1: How we started to use Vulkan in FFmpeg: a plea from the archival community, 2024: "FFv1 is too slow to encode, write CUDA or WHATEVER, just save us!". Slide 7, part 2: FFv1 is an IETF-standardized video codec designed for lossless compression. Supports RGB, YUV, Bayer, Floating point data (16 and 32 bits), CRCs. Adopted by the Library of Congress, amongst many other archival organizations. Slide 7, part 3: Luckily, I'm a Khronos member. Slide 7, part 3: I thought about it for a bit: "I mean, I guess I can parallelize Golomb mode using this neat algorithm I stayed up all night designing..." Slide 7, part 4: "No, we want range coding". "It's impossible". Slide 8: After some convincing, it turns out that FFv1 allows up to 1024 slices per frame. With an easily parallelizable prediction step . Slide 9: The range coder is not able to be parallelized. But we can use helper invocations to parallelize the RCT step, and load probabilities and prediction results. Slide 9.1: Why hybrid decoders never work: roundtrip is too expensive. Coefficients are usually larger than 8-bits, so they cost more to upload than the finished decode. dav1d attempted to move as many post-processing steps onto the GPU, but failed to gain over pure CPU decoding, even on low power devices. Slide 9.5 (sorry, I have to insert a slide, but I already numbered slides afterwards): Full FFv1 decoding process with GPUs. 32 invocations per slice, 1 invocation decodes, others do RCT and preload EC values. Packet data host mapped *onto* the GPU and read out via DMA. Slide 10, part 1: Benchmark results. 500fps on 4K 10bit. 2.5fps on an Alder Lake monster CPU to 30fps on a cheap 7900XTX on 6k5k16bit. Slide 10, part 2: this proved that even with codecs with no good parallelization methods, gain possibilities are enormous. Slide 10, part 3: But the key point is that decoding MUST happen on the GPU. Transferring hundreds of megabytes up and down PCI links has too much overhead, latency, and CPU load. Not to mention CPUs stall. Slide 10.5: The optimizations enabled FFv1 to move beyond basic archival. Slide 11: Now, back to ProRes. Better than FFv1, we can apply some advanced techniques and to proper parallel encoding AND decoding . Slide 12, part 1: Benchmark results. An old laptop integrated GPU can easily decode 2Gpbs 6k ProRes. Slide 13, part 1: So back to the camera. You'd like to record in RAW because cameras don't know what combination of red/white/blue is "white", and you're too poor for sets and good lighting. Slide 13, part 2: So you get the Panasonic Lumix S1ii. Its got a sensor equal in performance to a 50k Arri. Slide 13, part 3: Apple have you covered! At a price, despite already paying for it in the camera... and a limited operating system. Slide 14, part 1: Well, we have a binary, and we have Ghidra. Slide 14, part 2: Combine them and you get magic . Slide 15: So you dig into the code, tracing where every bit of the input ends up... . Slide 16, part 1: Then you get lost, but you write something good enough to see an image on screen. It doesn't look right . Slide 16, part 2: Eventually you get something right out... Slide 16, part 2: Ship it, we'll figure it out later, it's useful for users to monitor. Slide 17, part 1: You look around, and you notice there is a GPU version. Slide 17, part 2: But maybe the GPU version is cleaner... oh wait. Slide 18, part 1: CUDA kernel is encrypted. XOR with a 64-byte key somewhere. Good luck. Slide 18, part 2: Claude found it in a minute . Slide 19, part 1: It's not great. It's a CPU decoder bolted onto a GPU DCT , only for the data to travel back to the CPU. Slide 19: part 2: But, thanks to the decrypted kernel, we now have a full picture ! Slide 20: