On-device AI — professional-grade, fully private.

Add Text to GIF

Caption an animated GIF — meme text, subtitles or a watermark.

  • No upload
  • Meme caption bar
  • Timed per frame
  • Emoji supported

Drop a GIF here

GIF, animated WebP or APNG

Start from
Text layers
Your caption
Font
12%
92%
Placement
Placement

Or drag the caption on the preview.

0%

A plate behind the words. Set the solidity to 0 for none.

A caption bar makes the canvas taller and moves the picture, so nothing is covered.

Style
Alignment
8%
0%
100%
0°
Edges
Edges

Crisp adds only the colours you chose — better on photos and smaller files.

Show on frames
Show on frames

1 of 1 frames · 0.1s

Colours
256
Dithering

Off by default: a GIF's colours already came from a table, so dithering only adds noise.

How to add text to a GIF

Burn a caption into every frame of an animated GIF, with the animation's timing left exactly as it was.

  1. 1

    Add your GIF

    Drag an animated GIF onto the drop area, or click to browse. Animated WebP and APNG work too. The file is read on your device and never uploaded.

  2. 2

    Type your caption

    Type into the text box and the preview updates as you go. Long lines wrap on their own, and pressing Enter starts a new line where you want one.

  3. 3

    Place it and style it

    Pick one of nine positions or drag the caption anywhere on the picture. Choose a font, a size, an outline and a colour. For the classic meme look, switch the layer to a caption bar and it grows the canvas above or below the picture.

  4. 4

    Download the captioned GIF

    Press Add text, then Download. Every frame's delay, the total duration and the loop count come through unchanged, so the animation plays exactly as it did before.

Why add text to a GIF

An animated GIF plays without sound, without controls and usually without any context around it. Text is the only thing that can carry the joke, name the product, label the step in the tutorial or credit the source — and it has to be part of the picture, because a GIF gets copied out of its page and pasted somewhere else within minutes of being published. A caption written into the pixels travels with the file; a caption written in the surrounding HTML does not.

Up to eight text layers, four bundled fonts that look the same on every device, outlines and shadows, a caption bar that grows the canvas, and text timed to a range of frames.

Reads GIF, animated WebP and APNG. Writes a standard GIF89a with frame count, per-frame delays, total duration and loop count preserved exactly.

What a burned-in caption buys

  • It survives the copy. A GIF is dragged into a chat, re-uploaded to a forum and embedded in a slide deck. The only text that comes with it is the text inside the frames.
  • It works with the sound off. GIFs have no audio track at all, so a demo, a reaction or a tutorial step has to say what it means in writing or not at all.
  • It makes a reaction into a meme. The top-and-bottom caption is a format people recognise instantly, and it is the difference between a clip and a joke.
  • It credits and it brands. A small semi-transparent line in a corner survives every re-post, which is why watermarking a GIF is a caption problem rather than an image-editing one.

How text gets into an animation

A still image has one raster to draw on. An animated GIF has a stack of them, each potentially covering only part of the canvas, each with its own delay and its own transparency mask, and each stored as the difference from the frame before it. Text has to end up on every one of those frames, identically, or the animation develops a shimmer along the letter edges that no viewer can name but everyone can see.

What happens per frame

  • Compose. Every frame is drawn onto a canvas of the animation's logical size, honouring its offset, its disposal method and its transparency, so the caption lands on the picture a viewer would actually see.
  • Rasterise once. The caption is drawn into a single overlay, not once per frame. Every frame then receives literally the same pixels, so drift and flicker are impossible rather than merely unlikely.
  • Seed the palette. Your text colours and the gradient between fill and outline are put into the colour table by name, before the picture's own colours are derived from a histogram.
  • Re-optimise. Regions that did not change between frames are marked transparent again, which is where most of a GIF's size goes — and a caption that never moves costs almost nothing after the first frame.

Identical on the first frame, the middle frame and the last

Because the caption is one raster composited many times, the check is a strict one: on a 100-frame animation, the pixels that make up the letters are compared on frame 0, frame 50 and frame 99 and must match exactly. They do. A tool that draws the text per frame instead cannot promise this — a font that finished loading between frame 2 and frame 3, or a transform that rounded differently at a different offset, produces a caption that is subtly not the same shape, and GIF's frame differencing then re-sends it in full on every frame.

The palette problem, which is this tool's real subject

Every other GIF tool moves pixels that already existed. This one invents them, and that puts it straight into the format's oldest constraint: a GIF holds at most 256 colours for the whole animation. Anti-aliased letters are precisely what a full table cannot absorb. The soft edge of a letter is a gradient — from the fill colour to the outline, and from the outline into whatever the picture happens to be — and on photographic footage those intermediate values have nowhere to go. The quantiser snaps each one onto the nearest colour it already had, and the caption comes back chewed. That is why captioned GIFs from most online tools look worse than the same caption on a PNG.

Two answers, both measured

  • Seeded colours, always on. Your fill, outline, shadow and bar colours, plus six steps of the fill-to-outline gradient, are written into the palette by name before the picture's colours are derived. Before this existed, a caption specified as #ffffff came back with not one pure-white pixel in the output — nothing about which looks wrong.
  • Crisp mode, when you want it. Hard-edges the caption against both boundaries: the alpha against the picture, and the fill against the outline. The caption then contributes two colours to the table instead of sixty-four.
  • The measurement. Against a full-colour render of the same composite, on a photographic fixture with no free palette entries: Smooth 2.329 of 255 on the caption's pixels, Crisp 0.000. The picture behind the caption improved as well, from 3.594 to 0.168.
  • What was tried and dropped. Re-weighting the colour histogram toward the caption's region was implemented, measured and removed: once the colours are seeded by name it bought nothing on the letters and cost the picture accuracy at every palette size tested.

Why dithering is off by default here

Every sibling tool in this section dithers by default and this one does not, because the input is already an animation whose colours came out of a table of at most 256. Dithering re-approximates pixels that were already exact, across the whole picture, to buy nothing. Measured on frame 0 against a full-colour render: background error is 0.000 of 255 with dithering off, and between 8.9 and 11.9 with ordered dithering on. Ordered and error-diffusion dithering stay available for a true-colour source — an APNG or an animated WebP — where there is a real gradient to spread.

A caption bar, or text on the picture

There are two ways to put words on an animation and they are not interchangeable. Overlaying puts the text on top of the picture, which costs nothing in size or shape and covers whatever is underneath. A caption bar adds a band above or below and moves the picture down or up to make room — so the canvas gets taller and nothing is hidden. That second one is the classic meme format, and it is the one most in-browser tools do not offer at all, because growing the canvas means every frame has to be re-composited at a new size.

Which to use

  • Overlay, for reaction GIFs. The picture is the joke and the words are a comment on it. Use an outline so the text survives whatever moves underneath it.
  • Bar above, for the meme format. The set-up line reads before the picture. The canvas grows by exactly the height the wrapped text needs plus padding — two lines make a taller bar than one, automatically.
  • Bar below, for subtitles and credits. A black band under the picture reads as a subtitle plate and never covers the action. Bars above and below can be used together.
  • A transparent bar. Grows the canvas without inventing a background, which is what a sticker with a caption underneath needs.

The bar's height is measured, not guessed

The band is exactly as tall as the wrapped text needs plus padding, worked out before a single frame is drawn — because the encoder is given one width and height for the whole animation and refuses a frame that arrives at another size. Type a longer caption and the bar grows a line; shorten it and the canvas shrinks back. A 200×120 GIF with a two-line bar above becomes 200×172, and the picture is pushed down rather than covered.

Outline, shadow, and text you can actually read

A caption overlaid on a moving picture has no fixed background. White text is invisible over a bright frame and black text disappears into a dark one, and because the picture is moving, a colour that works on frame 1 fails on frame 40. The outline is the answer, and it is not a decoration: it puts a constant border between the letter and whatever is behind it, so the contrast is fixed by you rather than by the footage.

The legibility controls

  • Outline width. Measured as a percentage of the font size, so it scales with the text rather than turning to a hairline on a big caption. The stroke is painted at double the requested width and the fill goes on top, which puts the whole outline outside the letter instead of eating half of it.
  • Shadow. A soft offset drop shadow, for a subtler separation than a hard outline gives. It is drawn with the outline pass only, so a multi-line caption does not accumulate a second copy of it through the semi-transparent edges.
  • Opacity. For watermarks and credits, where the point is to be present without competing with the picture. It applies to the whole layer, outline included.
  • A background plate. A filled box behind the words, at whatever solidity you choose. It fixes the contrast completely rather than improving it, and on a small palette it is the cheapest of the four: one flat colour instead of an anti-aliased edge around every letter.

The side effect nobody mentions

An outline is also what makes anti-aliased text survive a small palette. Without one, the soft edge of every letter blends into the picture, so the intermediate colours depend on the footage and cannot be predicted. With one, the interior edge blends from the fill into the outline — a short, known gradient that can be written into the colour table in advance. That is why the six seeded ramp steps exist, and why an outlined caption comes out cleaner than an outline-free one on the same GIF.

Text that appears and disappears

A caption does not have to be on the whole animation. Each layer has a frame range, so it can run from frame 10 to frame 30 and be absent elsewhere. Two lines of dialogue become two layers with different ranges; a punchline can be held back until the moment it lands; a tutorial can label each step as it happens. The panel shows how many frames a layer covers and how long that is in seconds, because a frame number means nothing without the delays beside it.

How it stays cheap

  • One overlay per distinct set of layers. A whole animation with one always-on caption rasterises exactly one overlay. Two timed layers with different ranges produce three or four, not one per frame.
  • The bar is not timed. A caption bar changed the canvas size, so the band belongs to the whole animation and only the words come and go. Blinking the band would read as a rendering fault rather than as text appearing.
  • Ranges are clamped to the animation. A range that runs past the last frame stops at the last frame instead of silently producing a layer that never draws.

Fonts that render the same on every device

This is the quiet failure in nearly every captioning tool. A font is named, the browser or the server resolves it, and if it is missing something else is substituted — with no error and no visible sign. Measured in this repo on a different tool: two entries in a fifteen-font picker produced text metrics identical to the fallback family, meaning the canvas had been rendering something else the whole time while the dropdown claimed otherwise. Impact, the meme font, only appears to work because the machine testing it was usually Windows; it is absent on most Android, iOS and Linux devices.

The four bundled faces

  • Anton. The open substitute for Impact, and the default. Heavy condensed capitals — the meme voice. Bundled rather than named, so every device gets the real face.
  • Montserrat. A geometric sans at bold weight. The neutral choice for product demos, labels and anything that should not look like a joke.
  • Oswald. Condensed, so long captions fit across a narrow GIF without shrinking to nothing. The subtitle face.
  • Caveat. A handwriting face, for annotations and arrows-and-notes explanations where a mechanical font reads as too formal.

How the tool proves the font took

Each face is registered from its own font file before a single letter is measured — which matters because measuring with an unregistered face returns the fallback's widths, and the layout computed from them is then wrong for the face that eventually draws: lines wrap in the wrong places and a caption bar comes out the wrong height. After registration the tool asks whether the family really resolves, and reports on screen if it did not. It never quietly draws in a system font and calls it Anton.

Emoji, accents and other scripts

The bundled fonts cover Latin and Latin Extended, which is every character in English, Spanish, French, German, Portuguese and Indonesian. Anything outside that falls through to your own device's fonts, per character, so the letters around it stay in the chosen face. That is a deliberate trade: bundling a font that covers every script would mean shipping tens of megabytes to caption a 200 KB GIF.

What that means in practice

  • Emoji render in colour. The fallback chain names the platform emoji families explicitly, because a generic sans family resolves to a text font and a text font draws an empty box. Measured: a caption with two emoji introduces 980 distinct colours to the overlay against 30 for the same caption without them — they really are colour artwork, not outlines.
  • Accented Latin is exact. Both the Latin and Latin Extended subsets of each face are registered, so a French or Portuguese caption does not change font halfway through a word.
  • Other scripts use your device's fonts. Cyrillic, Greek, Arabic, Hebrew, Thai and CJK all render if your device has a font for them, which every modern device does. They will look like your system's face rather than like ours.

The meme layout, in one press

Top-and-bottom capitals in a heavy condensed face with a black outline is a format, not a style choice — it is recognisable enough that getting it slightly wrong reads as amateurish. The Meme preset sets both layers at once: Anton, uppercase, white fill, a black outline at the width the original format used, one layer anchored top and one bottom, sized as a proportion of the picture so it looks right on any GIF.

Getting it right

  • Two layers, not two lines. The set-up and the punchline are separate layers so they can be positioned and timed independently. A single layer with a newline in it puts both lines in the same place.
  • Capitals, deliberately. The uppercase toggle folds the text as it is drawn rather than asking you to type in caps, so the same caption can be switched between formats without retyping.
  • Bars for the caption format. The other recognisable meme shape puts the set-up on a white band above the picture instead of over it. Switch the placement and the canvas grows to fit.

Writing a caption that reads

A GIF loops every two or three seconds and a reader gets one pass at the words before the picture starts again. That is a much harder constraint than a still image, and it is the reason most captioned GIFs fail: the text is too long to finish before the loop, or too small to read at the size it will be seen, or placed exactly where the motion is.

By material

  • Reaction GIFs. Six words or fewer, large, outlined, near an edge. If it does not fit on two lines it is not a reaction caption.
  • Product and UI demos. Montserrat, a caption bar below rather than an overlay, and a frame range per step so each label appears with the action it describes.
  • Tutorials and how-tos. Number the steps and time each layer to its step. A permanent caption listing four steps is unreadable while a fifth thing is moving.
  • Watermarks and credits. Small, low opacity, a corner anchor, and no outline — a watermark that competes with the picture gets cropped out by whoever re-posts it.

Size is a percentage on purpose

Every measurement here — the font size, the outline width, the shadow, the wrap width — is a percentage of the picture rather than a pixel count. A 24-pixel caption is unreadable on a 640-pixel GIF and covers a 90-pixel one, so a pixel size is a setting you would have to re-choose for every file. As a percentage, one set of settings produces a consistent-looking caption across a whole batch of differently sized animations.

Where a captioned GIF is the answer

Captioning is not one job. A meme, a product demo, a bug report and a watermark want different fonts, different placements and different timing, and the settings that make one of them good make the others worse.

By purpose

  • Memes and reactions. The format the top-and-bottom layout exists for. Recognisable at a glance and readable in one loop.
  • Product demos in a README. A screen recording with a labelled step is documentation; the same recording without labels is decoration.
  • Support and bug reports. An arrow-and-note annotation on a recording of the problem saves a paragraph of description.
  • Social posts with the sound off. Every autoplaying animation on a feed is silent, so anything it needs to say has to be written on it.
  • Stickers and emoji. A short word on a transparent sticker, with a transparent caption bar so the words sit outside the artwork.

Which animations this opens, and what it writes

The input list is closed and enforced at the file dialog rather than at the engine, so a file this tool cannot open is refused in milliseconds with a route to the tool that can — instead of being accepted and failing later behind a spinner.

  • GIF in. Animated or single-frame, interlaced or not, with a global palette or a different local palette on every frame. Sub-rectangle frames at an offset are composited onto the full canvas before anything is drawn on them.
  • Animated WebP and APNG in. Both are read by our own demuxers on every browser rather than by a native decoder, so frames and timing are identical everywhere.
  • GIF out, always. A standard GIF89a with a global colour table, per-frame delays, disposal methods and a Netscape loop extension.
  • Anything else is refused with a route. A video is sent to Video to GIF, a still image to the Meme Generator, which does the same job for a picture that does not move.

What the tool reads

GIF is decoded natively where the browser has a decoder for it and by our own reader everywhere else, so an animation opens the same way in every browser rather than arriving as its first frame. Animated WebP and APNG always take our own reader, which is what keeps frame counts and timing identical across engines.

Size and batch limits

  • Free. One file up to 50 MB and up to 300 frames, with three text layers and the full 256-colour palette. The free plan is not a reduced-quality caption — it is a single-file one.
  • Pro. Up to 20 files at once, downloaded as one ZIP, and eight text layers. Batch is what Pro buys here, because captioning a set of clips with the same credit line one at a time is the actual chore.
  • The limit both plans share. Every frame is expanded to full colour while it is being worked on, so a memory budget bounds the job before it starts rather than letting the tab stop responding while the progress ring keeps turning.

How a caption is drawn on your device

Nothing is uploaded. The font is registered from bytes that came with the page, the caption is measured and laid out, one overlay is rasterised, the file is decoded frame by frame, each frame is composed with the overlay, a palette is built with your colours already in it, unchanged regions are marked transparent again, and the result is written out as a standard GIF89a. All of it runs in a background thread so the page stays responsive, and the colour mapping runs on your graphics hardware where it is available.

Why the same file captions identically every time

  • Fonts that came with the page. No network request resolves a font at run time, so a caption cannot depend on what a font server was doing that minute, or on which fonts a device happens to have installed.
  • A deterministic quantiser. Median cut, not the randomised neural quantiser most GIF libraries use, so the same input with the same settings always produces byte-identical output.
  • Identical results on every path. The graphics-accelerated and processor paths are required to produce the same bytes, not merely similar ones — a path choosing a different palette entry would show as colour flicker between frames.

What a captioning tool must not change

Four things have to come out exactly as they went in: the number of frames, each frame's own delay, the total playback time and the loop count. A tool that re-times an animation while captioning it produces a perfectly valid GIF that simply plays at the wrong speed, and nothing about the file looks wrong. The one dimension that may change is the canvas height, and only when you asked for a caption bar — which is a request rather than a side effect.

The caption is rasterised once and composited onto every frame, and the colour mapping is hardware accelerated where your device supports it, with the accelerated and fallback paths verified to produce byte-identical output. Measured on the reference device: a 100-frame 480×270 animation captioned in 284 ms. It all runs on your device, not our servers.

Frequently asked questions

Will the caption look the same on my friend's phone as it does here?

Yes, because the fonts travel with the tool. Most captioning tools name a font and let each device resolve it — which quietly renders your caption in Impact on Windows, in something else on Android, and in a system default on Linux, with no warning that a substitution happened. The four fonts here (Anton, Montserrat, Oswald and Caveat) are bundled as font files and registered before a single letter is measured, so the line breaks, the letter shapes and the caption bar's height are identical on every machine. The tool tells you on screen if a font ever fails to register rather than silently drawing in something else.

Why does text on a GIF often come out with rough, speckled edges?

Because a GIF holds at most 256 colours for the whole animation, and anti-aliased letters need a lot of them. The soft edge of a letter is a gradient between the text colour and whatever is behind it, and on a photographic GIF the table is already full — so the quantiser snaps each of those intermediate greys onto whatever colour it happened to have, and the letter edges break up. This tool answers it two ways. Your text colours, plus six steps of the fill-to-outline gradient, are put into the palette by name before the picture's colours are derived. And Crisp mode hard-edges the letters so they contribute exactly the two colours you chose and nothing between them: measured against a full-colour render of the same composite, that is 0.000 of 255 error on the caption's pixels.

Can I make the caption appear only for part of the animation?

Yes. Every text layer has a frame range, so a caption can run from frame 10 to frame 30 and be absent everywhere else — which is how you subtitle a two-line exchange or make a punchline land. The tool shows how many of the animation's frames the layer covers and how long that is in seconds. If the layer is a caption bar, the band itself stays for the whole animation, because it changed the canvas size; only the words come and go. A blinking band would look like a rendering fault.

How do I get the classic white-bar meme look?

Switch a text layer's placement from Overlay to Caption bar above. That grows the canvas: a 200×120 GIF with a two-line bar becomes 200×172, with the picture pushed down and the words on a solid band rather than on top of the image. Nothing is covered. You can have a bar above and a bar below at once, and the bar colour is yours — white for the classic look, black for a subtitle plate. The Meme preset sets all of this in one press, along with heavy Anton capitals and a black outline.

Does adding text change the animation's speed or looping?

No. Each frame keeps its own delay, the total playback time is identical to the millisecond, and the loop count comes through as it was — a GIF authored to play three times still plays three times. This is worth checking on any tool you use, because re-timing an animation while captioning it produces a perfectly valid GIF that simply plays at the wrong speed, and nothing about the file looks wrong. It is asserted here on every run.

Will the caption drift or flicker between frames?

It cannot, because it is only drawn once. The caption is rasterised into a single overlay and that same overlay is composited onto every frame, so every frame receives byte-identical caption pixels. This matters for more than tidiness: GIF stores each frame as a patch over the one before it, so text that moved by a fraction of a pixel between frames would be re-sent on every frame and multiply the file size. Measured on a 100-frame animation, the caption's pixels are identical on the first frame, the middle frame and the last.

Can I use emoji?

Yes, and most tools cannot. A server-side renderer with a handful of Latin fonts installed has no glyph for an emoji and writes an empty box. Here the text is drawn on your device, so the fallback chain reaches your own emoji font: type an emoji and it renders in colour. The artwork is your platform's — an emoji looks like Apple's on an iPhone and Microsoft's on Windows, exactly as it does everywhere else on the web. The same applies to scripts outside the bundled fonts' coverage: accented Latin is bundled, and Cyrillic, Greek, Arabic or CJK fall through to whatever your device has.

Should I use Smooth or Crisp text?

Smooth is the default and looks better at large sizes, where the anti-aliased curve of a letter is what makes it read as a letter. Crisp is better on small GIFs, on busy photographic footage, and any time file size matters: it gives the caption exactly the colours you named, which measured 0.000 of 255 error against a full-colour render where Smooth measured 2.329, and it left the picture behind it more accurate too, because the palette entries the gradient was consuming go back to the photograph. Try both on the preview — the difference is easiest to judge at the size the GIF will actually be seen.

How many captions can I add at once?

Three on the free plan and eight on Pro. Three is enough for the two things people actually do — a top-and-bottom meme, or a caption plus a credit line — and each layer carries its own font, size, colour, position, rotation and frame range, so they do not have to look alike. Layers are drawn in order, so a later layer sits on top of an earlier one where they overlap.

Can I caption a transparent GIF without getting a black box?

Yes. A see-through GIF composited onto a blank canvas becomes fully transparent black, which then maps to the nearest palette entry to black — which is how a captioned sticker ends up with a dark rectangle behind it on other tools. Here every frame is checked for transparency before anything is encoded, and if any pixel of any frame is see-through the transparency is preserved through the caption. A caption bar can be transparent too, which puts the words outside the picture without inventing a background.

Why is my captioned GIF bigger than the original?

Because text is new information. GIF compresses by storing each frame as the difference from the one before, and a static caption compresses well — but the letters themselves are pixels that were not there, and their colours take palette entries away from the picture. Two things help: Crisp mode, which typically produces a smaller file than Smooth because it adds two colours instead of sixty-four; and running the result through the GIF Optimizer or the GIF Compressor afterwards, which is what those tools are for.

Can I caption several GIFs at once?

On Pro, yes — up to 20 files in one go, downloaded together as a single ZIP. All of them get the same layers, and because every size here is a percentage of the picture rather than a pixel count, the same settings produce a consistent-looking caption on files of different dimensions. The free plan handles one file at a time, at full quality and with the full 256-colour palette.