What you get
Pick PNG or JPG, choose how often to grab a frame, and the converter writes out numbered stills — frame_0001, frame_0002, and so on — then packs them into a single ZIP so you download one file instead of several hundred.
The numbering is zero-padded deliberately. Sequences named frame_1, frame_2 … frame_10 sort wrongly in almost every file browser and in ffmpeg's own pattern input, which puts frame_10 immediately after frame_1. Four digits of padding keeps the order correct everywhere, which matters if the images are going back into a video later.
The ZIP itself is stored rather than compressed. PNG and JPG are already compressed formats, so running them through deflate again costs time and saves almost nothing — on a few thousand frames that difference is minutes.
Choosing how many frames
You can also set an output width, and the height follows automatically to preserve the aspect ratio. Scaling uses Lanczos, which holds onto detail better than the default when you're reducing size — worth having when the frames are going to be looked at rather than processed.
- Every frame — a true frame-by-frame dump. A 30-second 60fps clip gives you 1,800 images, so check the length before you start.
- A fixed rate — one image per second, two, ten. Good for contact sheets, thumbnails, or skimming a long recording for the moment you need.
- Every Nth frame — keep every 5th or every 10th. Thins out a sequence while staying tied to the original frame timing rather than resampling it.
Getting one specific frame
Set a trim range before extracting and you only get frames from that window. The trim is applied on the output side, which makes it frame-accurate rather than snapping to the nearest keyframe — so asking for 12.4 to 12.5 seconds genuinely gives you the frames in that tenth of a second, not whatever keyframe happened to be nearby.
That's the quickest route to a single still: narrow the trim to the moment you want, extract every frame, and pick from the handful you get.
PNG or JPG
PNG is lossless. Every frame is exactly what the decoder produced, which is what you want if the images are going into a compositor, being analysed, or going back out as video. The files are several times larger.
JPG is lossy but far smaller, with an adjustable quality setting. Sensible for contact sheets, reference stills, or anything headed straight for a web page. Not sensible if the frames are an intermediate step on the way to another encode, because you'd be stacking JPG loss underneath whatever the next encoder does.
Colour is converted, not guessed
Video stores colour as YUV; images are RGB. That conversion needs to know which matrix and range the source used, and getting it wrong is what makes extracted frames come out subtly dark, washed out, or with a slight colour cast — a mistake easy to miss on one frame and obvious across a sequence.
The converter probes the source's actual matrix and range and uses those values for the conversion rather than assuming bt709 limited. An HDR source is handled on the way out too — see converting HDR to SDR for what that involves.
Interlaced sources
Footage from DV tape, DVDs, or broadcast capture is often interlaced: each frame holds two fields captured at different moments. Extract frames from it untouched and you get combing — horizontal teeth along anything that moved. There's a deinterlace option that resolves the fields into whole progressive frames first, which is what you want on anything from a camcorder or a TV capture.
Why in the browser
Frame extraction is the operation where uploading makes least sense. The output is bulkier than the input — often by a wide margin — so a server-based tool has to receive your video, write out thousands of files, and send them all back. Doing it locally skips both transfers entirely, and the video never leaves your device. There's no size cap beyond what your own machine can hold.