Multimodal input

Attach images, PDFs, audio, and video to any prompt. Paste a screenshot with Ctrl-V, or just mention a file's path.

You don't have to describe a problem in words. A screenshot of a broken layout, a PDF spec, or a recording of a bug can ride along with your prompt. This works on every input surface: the interactive deck, stella run, and stella chat.

Ways to attach

Paste an image with Ctrl-V. In the deck, Ctrl-V reads your clipboard directly. If you copied an image, such as a screenshot or a figure, it's saved and attached to your prompt as a chip in the composer. Plain text still pastes normally, so Ctrl-V is safe to use for every paste.

Terminals can't normally receive a pasted image. A regular Cmd-V of a screenshot arrives as nothing. That's why this explicit shortcut exists. Pasted images are saved as PNG files in .stella/attachments/ in your workspace, so your session has a stable file to replay later. The store keeps your most recent 100 pastes and removes older ones.

Mention a file's path. Name a media file in your prompt and stella attaches the file itself:

stella run "the header overlaps the nav on mobile — see shots/header.png"

@-mention a file. An @-mention is a stronger signal than a plain path. Media files get attached, and text-like files (source code, markdown, JSON) get inlined as text. A plain path-shaped word never gets inlined as text on its own — prose mentions paths all the time, and inlining every one would flood the model with context. If an @-mention doesn't match a real file, stella leaves it as plain text.

stella run "implement the retry policy described in @docs/rfc-042.pdf"

What each kind becomes

You attachThe model receives
Image (image/*)A native image block. The model sees the pixels.
PDFA native document block
Audio, videoNative media, where the provider supports it (see below)
Text-like file (via @)Its contents, inlined as text
Any other binaryA text note describing what was attached

stella has no size limit of its own. Attachments live on disk and are only converted to base64 when a request is built, so only your provider's own limits apply. stella does cap you at 32 attachments per prompt, to stop a runaway request.

What each provider can see

Providers differ in what they can look at, and stella tracks these differences instead of assuming every provider works the same way:

ProviderImagesPDFsAudioVideo
anthropicyesyes——
openaiyesyes——
gemini, vertexyesyesyesyes
bedrockyesyes—yes
OpenAI-compatible (zai, openrouter, xai, deepseek, local)vision models only———

On the OpenAI-compatible path, image support depends on the model. GLM vision models (the v in glm-4.5v) can take images; plain GLM text models can't. For other vendors reached through a gateway, stella assumes the model can see images rather than silently dropping them.

Unsupported attachments

If a model can't handle an attachment, the turn still works. Sending an image to a text-only model would normally fail the whole request and lose your prompt. Instead, stella swaps the attachment for a short text note that tells the model what was attached: the filename, the kind, and that it can't view it. The model can then say so, or open the file with a tool when that makes sense, such as reading a text-extractable PDF, instead of guessing at contents it never saw.

In practice, this means you can paste a screenshot without checking which model is running first. A vision model sees it. Anything else knows what it missed, and tells you.

Right now, video works natively on Gemini, Vertex, and Bedrock. Everywhere else, it degrades to the text note.