FFmpeg on serverless: CPU vs GPU, cold vs warm, cost per job
We needed real numbers before choosing where KinoPipe renders: which worker class, CPU or GPU, and what a job actually costs once cold starts and failed bursts are counted. This is the full report from the August 2026 campaign, FFmpeg 9.0.1 on RunPod Serverless and Google Cloud Run, with every figure we used to pick a provider.
Measured 25 August 2026. Published 26 August 2026. 325 RunPod runs and 95 Cloud Run runs, 60 s fixtures, at least 5 repetitions per comparison.
8.2s
p50, 1080p H.264 fast, RunPod 16 GB GPU pool
8.9s
p50, same recipe, RunPod 8 vCPU CPU at 16 threads
$0.63
compute per 1,000 jobs, RunPod 4 vCPU CPU
100%
burst success at concurrency 4 on every RunPod class
TL;DR
- A 16 GB GPU pool on RunPod renders a 60 s 1080p H.264 fast transcode in 8.2s p50 with tuned NVENC. An 8 vCPU CPU worker does it in 8.9s at 16 threads, a 4 vCPU worker in 13.5s. For plain 1080p, CPU is within a second or two of GPU.
- Price separates the classes more than speed: $0.168/h for 4 vCPU, $0.576/h for the GPU pool, $1.224/h for an A40 that was not faster than the cheaper GPU pool on any recipe we run.
- Cold starts are real: 18.1s and 11.4s for the first CPU encode on 4 and 8 vCPU, 23.9s for the first GPU encode after an image roll, against 6.8 to 8.7 s warm.
- Cloud Run CPU cost about 2.5x the active hourly price and 2.2x the H.264 latency of RunPod CPU. Cloud Run L4 is close on 1080p (9.6s vs 8.2s) but 26.6s vs 11.0s on 4K to 1080p, at about 1.8x the hourly price.
- Bursts: every RunPod class finished 20 jobs at concurrency 4 with 100% success. Cloud Run 8 vCPU finished 65% at concurrency 4 (the rest failed with "no available instance") and 100% at concurrency 3.
What we measured, and how
Both campaigns ran the same FFmpeg 9.0.1 build, the same 60-second fixtures stored in Cloudflare R2 (a 1080p source, a 4K source, a 720p clip for concatenation and picture-in-picture, a captions file and an alpha watermark) and the same recipe matrix, which is the set of constrained operations KinoPipe exposes as typed tools: H.264 fast and balanced transcodes, 4K downscale, portrait 9:16 crop, subtitle burn-in, watermark, picture-in-picture, two-clip concatenation, VP9, a 4 s GIF and MP3 extraction. GPU runs target H.264 through NVENC. VP9, GIF and audio stay on CPU because the current pipeline does not accelerate those codecs on NVIDIA hardware.
The RunPod campaign retained 325 successful runs and no failed measured run, with separate cold, warm, burst and tuning phases and at least 5 repetitions per comparison. Its warm matrix covered eleven CPU recipes and eight NVENC-compatible GPU recipes. RunPod latency is end to end, delayTime + executionTime, with client polling delay stored separately. The Cloud Run campaign retained 95 successful warm runs, one warm instance of each type, 5 repetitions per provider and scenario; latency there is direct HTTPS end to end including the R2 input download and output upload.
Every run records end-to-end latency, worker total time and FFmpeg-only time, download, probe and upload time, FFmpeg speed versus real time, input and output bytes, the exact CPU model, logical CPU count, RAM, FFmpeg build and visible NVIDIA GPU, and an estimated job cost from the configured hourly rate. Cost figures below are estimates from those rates: RunPod billing data was still lagging the final campaign when the endpoints were deleted.
| Class | Provider | Billing | Rate |
|---|---|---|---|
| CPU 4 vCPU / 8 GB | RunPod Serverless | per second, workersMin 0 | $0.168/h |
| CPU 8 vCPU / 16 GB | RunPod Serverless | per second, workersMin 0 | $0.336/h |
| 16 GB Ampere pool (measured: RTX A4000) | RunPod Serverless | per second, workersMin 0 | $0.576/h |
| A40 48 GB | RunPod Serverless | per second, workersMin 0 | $1.224/h |
| Cloud Run CPU 8 vCPU / 16 GiB | Google Cloud Run | request-based, active time, plus $0.40 per million requests | $0.8352/h active |
| Cloud Run 1x NVIDIA L4 / 4 vCPU / 16 GiB | Google Cloud Run | instance-based, 60 s minimum, no zonal redundancy | $1.04652/h |
The RunPod 16 GB endpoint requested the Ampere category; the worker actually measured was an RTX A4000. RunPod Serverless does not guarantee a specific GPU model inside a generic category, which matters when you read the GPU rows below.
Warm results: p50 per class
The three headline recipes on RunPod, warm workers, 60 s inputs. "Tuned" means the NVENC CQ and thread settings adopted at the end of the campaign, described further down.
| Class | Rate | 1080 fast | 1080 balanced | 4K to 1080 | Best use |
|---|---|---|---|---|---|
| CPU 4 vCPU / 8 GB | $0.168/h | 13.5s | 21.9s | 28.5s | Lowest cost |
| CPU 8 vCPU / 16 GB | $0.336/h | 10.0s auto threads; 8.9s at 16 threads | 16.2s | 20.3s | CPU latency |
| 16 GB Ampere pool (measured: RTX A4000) | $0.576/h | 8.2s tuned NVENC | 9.3s | 11.0s | H.264 throughput |
| A40 48 GB | $1.224/h | 8.2s | 11.5s | 12.2s | No current advantage |
Two things stand out. The A40 at $1.224/h matched the cheaper GPU pool on 1080 fast and was slower on the other two recipes, so it buys nothing for these workloads. And the 8 vCPU CPU worker at 16 threads lands within 0.7 s of the GPU pool on 1080 fast; the GPU advantage only becomes decisive on the balanced preset and on 4K input.
Cost per job, derived
Hourly rates are hard to compare across providers with different billing bases, so here is the same table as compute cost for one thousand 60 s 1080p fast jobs, computed as rate multiplied by the measured p50 and nothing else. No storage, no egress, no queue wait, no minimum billing window.
| Class | Rate | p50 used | Per 1,000 jobs |
|---|---|---|---|
| CPU 4 vCPU / 8 GB | $0.168/h | 13.5s | $0.63 |
| CPU 8 vCPU / 16 GB, 16 threads | $0.336/h | 8.9s | $0.83 |
| 16 GB Ampere pool (measured: RTX A4000) | $0.576/h | 8.2s | $1.31 |
| A40 48 GB | $1.224/h | 8.2s | $2.79 |
| Cloud Run CPU 8 vCPU / 16 GiB, 16 threads | $0.8352/h active | 19.2s | $4.45 |
| Cloud Run 1x NVIDIA L4 / 4 vCPU / 16 GiB | $1.04652/h | 9.6s | $2.79 ($17.44 if each job wakes an instance) |
Cloud Run L4 is instance-based with a 60 s minimum, so an isolated job that wakes an instance is billed for at least 60 s. Cloud Run CPU adds $0.40 per million requests, negligible at this scale, and an idle minimum instance on that shape costs about $0.216/h before free-tier discounts. These are compute figures for provider selection, not KinoPipe's credit pricing, which also covers storage and delivery.
Cold starts
With workersMin: 0, the first request pays for the container start. On RunPod CPU the cold H.264 fast p50 was 18.1s on 4 vCPU and 11.4s on 8 vCPU, against 13.5s and 10.0s warm. The first encode after rolling the final GPU image took 23.9s, then 6.8 to 8.7 s for the following warm requests.
A fresh Cloud Run revision returned 13.5s on CPU and 10.9s on the L4 for the first request, but container timestamps show Cloud Run had started both workers about 19 seconds earlier as part of the revision rollout. We keep those as post-deployment, pre-warmed measurements, not as cold-start proof. A strict scale-to-zero cold campaign has to wait until the platform actually terminates the idle instances, and we have not run it yet.
Bursts and scale-out
Twenty jobs submitted at once, which is the shape of an agent fanning out a batch. Latency includes real scale-out where noted.
| Class and setup | Jobs | Concurrency | p50 | p95 | Success |
|---|---|---|---|---|---|
| CPU 4 vCPU, four ready workers | 20 | 4 | 14.4s | 17.1s | 100% |
| CPU 8 vCPU, four ready workers, 16 threads | 20 | 4 | 10.7s | 15.7s | 100% |
| 16 GB Ampere, scale one to four | 20 | 4 | 11.0s | 20.0s | 100% |
The GPU p95 of 20.0s includes scaling from one warm worker to four. One operational lesson: RunPod configuration PATCHes can report the new worker maximum before the data plane accepts requests, so control-plane propagation must be excluded from measured runs.
| Setup | Jobs | Concurrency | p50 (succeeded) | p95 (succeeded) | Success |
|---|---|---|---|---|---|
| CPU 8 vCPU, scale from one warm instance | 20 | 4 | 15.0s | 22.0s | 65% |
| CPU 8 vCPU, min=4 requested | 20 | 4 | 13.4s | 22.1s | 80% |
| CPU 8 vCPU, scale from one warm instance | 20 | 3 | 12.7s | 16.0s | 100% |
| GPU L4, scale one to three | 20 | 3 | 14.4s | 23.8s | 100% |
Cloud Run request logs classify the CPU failures as "no available instance": the client received 5 HTTP 500s and 2 HTTP 429s in the scale-from-one run, and only 2 additional instances started during the burst. Quota was not the limit: europe-west1 exposed 200 vCPU and 400 GiB while the test needed 32 vCPU and 64 GiB at concurrency four. Requesting four minimum instances only raised success to 80%. Concurrency three completed reliably. Direct synchronous FFmpeg traffic to Cloud Run therefore needs idempotent retry with backoff for 429 and infrastructure 500 responses, or, better, an asynchronous queue in front of it.
Cloud Run, recipe by recipe
The full warm matrix on Cloud Run, one warm instance of each type. The speedup column is CPU p50 divided by L4 p50.
| Scenario | CPU p50 | CPU p95 | L4 p50 | L4 p95 | L4 speedup |
|---|---|---|---|---|---|
| H.264 1080p fast | 22.2s | 23.1s | 9.6s | 11.0s | 2.3x |
| H.264 1080p balanced | 37.1s | 39.6s | 10.0s | 11.5s | 3.7x |
| 4K to 1080p | 47.3s | 53.0s | 26.6s | 27.4s | 1.8x |
| Portrait 9:16 | 36.7s | 37.6s | 12.4s | 13.6s | 3.0x |
| Burn subtitles | 38.8s | 39.3s | 11.1s | 12.0s | 3.5x |
| Watermark | 36.8s | 45.6s | 9.9s | 11.7s | 3.7x |
| Picture-in-picture | 36.5s | 44.6s | 11.8s | 14.0s | 3.1x |
| Concatenate two clips | 47.4s | 50.9s | 14.2s | 14.8s | 3.3x |
| Scenario | p50 | p95 | Note |
|---|---|---|---|
| VP9 software | 90.6s | 163.8s | Large host-placement variance |
| 4 s GIF | 13.2s | 15.7s | Palette generation dominates |
| Extract MP3 | 1.8s | 2.5s | I/O-heavy, inexpensive |
Cloud Run CPU performance turned out to be host-dependent even with an identical advertised shape. The same VP9 recipe took 86.4 to 90.6 s on AMD EPYC placements and 139.6 to 169.8 s on a generic Intel Xeon 2.80 GHz placement. The five 16-thread VP9 runs all landed on EPYC 9B14 and were much tighter, 82.9 to 86.5 s.
The whole Cloud Run warm matrix cost about $1.33 before free tier, network, storage, build and invoice rounding: $0.509 of active CPU time plus $0.032 of idle minimum, and $0.793 for the L4 over the 45.46-minute wall clock. The recorder's own sum of request-duration compute was $0.664, which understates the deliberately warm L4 because it excludes the time spent waiting for CPU runs.
| Class | Hourly rate | 1080 fast p50 | Burst p50 / p95 | Burst success |
|---|---|---|---|---|
| RunPod CPU 8 vCPU / 16 GB, 16 threads | $0.336/h | 8.9s | 10.7s / 15.7s at c4 | 100% |
| Cloud Run CPU 8 vCPU / 16 GiB, 16 threads | $0.8352/h active | 19.2s | 12.7s / 16.0s at c3 | 100% at c3; 65% at c4 |
| RunPod 16 GB Ampere pool (measured: RTX A4000) | $0.576/h | 8.2s tuned | 11.0s / 20.0s at c4 | 100% |
| Cloud Run 1x NVIDIA L4 / 4 vCPU / 16 GiB | $1.04652/h | 9.6s | 14.4s / 23.8s at c3 | 100% |
FFmpeg settings the campaign changed
Benchmarks are only useful if they change the defaults. Three did.
VP9
Balanced VP9 now uses deadline=good, cpu-used=4, row multithreading and two tile columns. Explicit encoder threads stay optional because FFmpeg's default was as fast or faster on the tested 4 and 8 vCPU shapes.
| Class | Before | After | Change |
|---|---|---|---|
| CPU 4 vCPU | 245.1s | 69.1s | -72% |
| CPU 8 vCPU | 238.9s | 44.9s | -81% |
Quality moved from SSIM 0.98278 to 0.98254 at an unchanged PSNR of 28.98 dB. Output size grew from 67.2 to 73.2 MB, a deliberate 9% storage trade for the compute reduction.
CPU threads
For H.264 fast, forcing threads made the 4 vCPU worker slower, so it keeps FFmpeg's auto-selection. On 8 vCPU, 16 threads improved p50 from 10.0s to 8.9s. On the Cloud Run 8 vCPU service, 16 threads beat 8 on H.264 fast (22.2s to 19.2s, -13.6%) and on VP9 (90.6s to 84.8s, -6.4%). Going from 4 to 8 FFmpeg and filter threads on the 4 vCPU L4 shape made 4K slower, 26.6s to 28.3s, so that shape keeps 4.
NVENC constant quality
The original CQ mapping over-produced quality and bytes. The validated defaults are now CQ 34 for fast and CQ 29 for balanced; the quality preset keeps its previous mapping because it was not part of this tuning gate.
| Scenario | Before | Tuned | Reduction | SSIM | PSNR |
|---|---|---|---|---|---|
| 1080 fast | 51.1 MB | 24.4 MB | 52% | 0.98969 | 41.93 dB |
| 1080 balanced | 77.9 MB | 44 MB | 44% | 0.99466 | 44.55 dB |
| 4K to 1080 balanced | 64.7 MB | 30.9 MB | 52% | 0.99648 | 46.33 dB |
Smaller outputs also shorten the R2 upload and cut storage and egress. One build detail worth knowing: the GPU images pin nv-codec-headers 12.2. The 16 GB pool exposed NVENC API 12.2 while the A40 exposed 13.0, and building against 13.x broke the cheaper pool.
What we decided
RunPod Serverless is KinoPipe's single compute provider, with two endpoint classes: the 4 vCPU CPU shape as the cost-first default for ordinary H.264, VP9, GIF and audio, and the 16 GB Ampere-category GPU pool for 4K, composition-heavy H.264 and latency-sensitive bursts. The A40 is out: more than twice the price and not faster for these recipes. Both endpoints start at workersMin: 0; a warm worker only gets added when observed traffic makes its latency benefit worth the continuously billed idle time. R2 stays the shared input and output store. The tested Cloud Run L4 service is technically sound and remains a fallback.
The trade-offs are explicit. Scale-to-zero means the first job after idle pays the 18.1s cold start on CPU. The generic GPU category can hand you a different card than the one you benchmarked. And any provider needs an asynchronous queue in front of FFmpeg for bursts, which is why every KinoPipe edit is a job you poll or receive by webhook rather than a synchronous request. That is also what lets compound edits compile to a single FFmpeg pass instead of a chain of re-encodes.
Limitations and what we did not measure
- Cost per job is derived from configured rates and measured p50, not from invoices; RunPod billing data was still lagging when the endpoints were deleted.
- The 4K pipeline still decodes and scales in software on the GPU. CUDA and NPP filters are the next optimisation to test and could change the 4K rows.
- No strict scale-to-zero cold campaign on Cloud Run yet; its cold numbers above are pre-warmed by the rollout.
- One region and one day per provider. Cloud Run CPU placement varied between AMD EPYC and Intel Xeon hosts within the same shape, which is a variance source we cannot control.
- The RunPod GPU category did not deliver the requested model; results are for an RTX A4000, not a guaranteed Ampere SKU.
- SaladCloud and Koyeb were set up but not run: the organisation and images exist, no paid container was created.
- Quality (SSIM and PSNR) was checked on the tuned recipes only, not across the whole matrix.
How to reproduce
The harness lives in the KinoPipe repository under benchmark/. It runs the same constrained recipes on RunPod, SaladCloud, Koyeb and Cloud Run, stores fixtures and outputs in R2 and writes JSONL, CSV and Markdown comparisons. It does not provision provider resources, and external provider calls are disabled unless --execute is passed. The steps, in order:
npm --prefix benchmark install npm run bench:check # local setup npm run bench:check:r2 # one R2 bucket HEAD request npm run bench:fixtures -- --generate npm run bench:fixtures -- --upload cp benchmark/config.example.json benchmark/config.json # enable providers, fill current rates npm run bench:run -- --execute npm run bench:report npm run bench:quality # SSIM / PSNR on the outputs
If you rebuild it elsewhere, keep the three things that make results comparable: the same FFmpeg revision on every provider, the same fixtures, and cold, warm and burst measured as separate phases. Comparing advertised hourly prices without startup, idle timeout and provider queue behaviour tells you very little.
FAQ
Is a GPU always faster than a CPU for FFmpeg on serverless?
Not by much for plain 1080p H.264. The RunPod GPU pool reached 8.2s p50 and the 8 vCPU CPU worker 8.9s with 16 threads. The gap opens on composition-heavy recipes and on 4K input, where the GPU pool did 11.0s against 20.3s on 8 vCPU. VP9, GIF and audio stayed on CPU because the current pipeline does not accelerate them.
How much does one FFmpeg job cost on serverless?
Compute only, at the measured p50 for a 60 s 1080p fast transcode: about $0.63 per 1,000 jobs on a 4 vCPU RunPod CPU worker, $1.31 on the 16 GB GPU pool and $4.45 on Cloud Run 8 vCPU. Storage, egress, queue wait and provider minimum billing windows come on top.
What does a cold start cost you in latency?
On RunPod CPU the first H.264 fast encode took 18.1s on 4 vCPU and 11.4s on 8 vCPU, against 13.5s and 10.0s warm. The first GPU encode after rolling a new image took 23.9s, then 6.8 to 8.7 s warm.
Why did KinoPipe pick RunPod over Cloud Run?
Price and burst behaviour. Cloud Run CPU cost about 2.5x the active hourly price and 2.2x the H.264 latency of RunPod CPU, and its 8 vCPU service completed only 65% of a 20-job burst at concurrency four. The L4 service is close on 1080p but 26.6s against 11.0s on 4K, at about 1.8x the hourly price.
These numbers back the p50 on the FFmpeg as a service page and the credit model on pricing. The recipes measured here are the ones behind compress video, add subtitles and video to vertical.