AM alexandermayorov.com
Belgrade
All articles

>_MCP · agents · file uploads

How to upload files through MCP: images and documents via curl

Why an agent shouldn't pass files as MCP tool arguments and how to upload them without touching the model's context: a signed 5-minute link and curl.

You can upload a file through MCP, just not as a tool argument. The MCP server gives the agent a short-lived signed link, and the agent sends the file itself with a plain curl from its own terminal. The bytes bypass the model's context, so they cost no tokens.

I have a project where an agent works through an MCP server, and the project needs to accept files and images from that agent. When it came to images, the agent answered honestly: "You can't hand an image to the agent: a person puts the file in place, and the agent only inserts the link." Then it suggested a workflow where a person opens a browser, goes to the file storage and uploads the file by hand.

I already had a solution. Now the agent uploads files to the project on its own, with no browser and no human in the loop. Here's how it works and why the direct approach doesn't.

Why you can't pass a file as an MCP tool argument

MCP tool arguments are JSON, and JSON has no type for binary data. So an image has to be turned into a string, usually base64. And that string has to be written by the model, because the model is what composes the call's arguments.

  • The whole file goes through the context. The agent first reads the file, then copies it into the call character by character. A 300 KB image grows to roughly 400 KB of text in base64. That's more than a hundred thousand tokens for a single image.
  • The agent doesn't need the file's contents. Its job is to put the file in the right place and insert a link into the text. It doesn't have to read the image's bytes for that.
  • The model doesn't copy data, it generates it. Get one character wrong in a long base64 string, and the server receives a corrupted file.
  • Limits. A large PDF or a slide deck simply won't fit into a single call.

I first ran into this earlier, on a different task. The agent saved finished reports to the server through an MCP tool that took the entire file as an argument. It worked, but every report went through the context in full. The agent was spending tokens just to move a file from a folder to a server, a job any terminal command can handle.

The solution: MCP coordinates, the file travels over HTTP

That's when I came up with the idea of splitting the two flows. MCP carries only the control messages: where to upload, what happened, what to insert into the text. The bytes themselves the agent sends as a regular HTTP request with curl. The model writes one short command with a path to the file, and curl reads and sends the file without the model's involvement.

The idea isn't new. S3 presigned URLs work the same way: the server doesn't pass the file through itself, it gives the client a temporary upload link. The difference is that here the client is an AI agent with a terminal.

  1. 01Agentasks the MCP server for an upload link
  2. 02MCP serverreturns a 5-minute link and a ready-made curl command
  3. 03curlsends the file's bytes, bypassing the model's context
  4. 04Receiverchecks the files and replies with lines for the text

I run two variations of this pattern:

  • A link from MCP. The agent calls a tool, gets a one-time link and a ready-made curl command, and runs it. It fits when you don't know the files and their destinations in advance.
  • Upload by convention, MCP only confirms. Following instructions from a skill, the agent uploads the file to storage with curl, then calls an MCP tool that says "file uploaded". The MCP server and the storage live on the same server and build the file path by the same rule. The server checks that the file is there and not empty, then marks the task as done. It fits when you know in advance which file goes where.

In both cases the agent never reads the file's contents. The model sees only a short request and a short response, and the file travels over a separate channel.

How the agent uploads an image, step by step

Let's walk through the first variation with a generic example. Names and addresses are illustrative.

Step 1. The agent asks for a link

The agent calls an MCP tool and passes only the ID of the record the files belong to. In return it gets a link that lives for 5 minutes and works only for that record, plus a ready-made command:

{
  "ok": true,
  "upload_url": "https://example.com/upload/<key>",
  "expires_at": "10/09/2026 3:05 PM",
  "curl": "curl -sS -F \"file1=@/path/to/image.png\" \"https://example.com/upload/<key>\""
}

The tool description states plainly that bytes are not passed through MCP and that curl has to run where the files are. Without that sentence, the model sometimes tries to pass the file the old way.

Step 2. The agent runs curl

curl -sS -F "file1=@./Diagram.png" -F "file2=@./Report.pdf" "https://example.com/upload/<key>"

A single request can carry several files, each under its own field name. The command doesn't touch the model's context: it contains paths to the files, not their contents.

Step 3. The server stores the files and replies with ready-made lines

The server checks the link, puts the files into the record's folder and renames them right away: "Diagram.png" becomes diagram.png. The response contains lines the agent pastes into the text as is:

{
  "ok": true,
  "saved": [
    {"file": "diagram.png", "markdown": "![caption](files/diagram.png)"},
    {"file": "report.pdf", "markdown": "[report](files/report.pdf)"}
  ],
  "errors": [],
  "next": "Insert the lines from saved[].markdown into the text and save the record."
}

The next field isn't for a human. It's a hint telling the agent what to do next. The agent reads the curl response the same way it reads an MCP tool response.

Step 4. The agent saves the result

The agent inserts the lines in the right places in the text, adjusts the captions and saves the record with a regular MCP tool.

In the end, a person types one sentence: "upload the images from this folder and add them to the text". The agent does the rest.

What the receiver looks like on the server

A simplified skeleton in TypeScript for Node.js with Express and multer. The signing details are deliberately left out, you write those yourself:

import express, { type NextFunction, type Request, type Response } from "express";
import multer from "multer";
import { MAX_FILE_BYTES, MAX_FILES } from "./config";
import { storeFile, type SavedFile } from "./storage";

type Grant = { recordId: string };

export function makeUploadKey(recordId: string): string {
  ...
}

export function verifyUploadKey(key: string): Grant | null {
  ...
}

const app = express();
const upload = multer({
  storage: multer.memoryStorage(),
  limits: { fileSize: MAX_FILE_BYTES, files: MAX_FILES },
});

app.get("/upload/:key", (_req: Request, res: Response) => {
  res.status(405).json({
    ok: false,
    message: 'POST multipart/form-data only: curl -F "[email protected]" <this link>',
  });
});

app.post(
  "/upload/:key",
  (req: Request, res: Response, next: NextFunction) => {
    const grant = verifyUploadKey(req.params.key);
    if (grant === null) {
      res.status(401).json({ ok: false, message: "The link is invalid or expired. Get a new one through MCP." });
      return;
    }
    res.locals.grant = grant;
    next();
  },
  upload.any(),
  async (req: Request, res: Response) => {
    const grant = res.locals.grant as Grant;
    const files = (req.files ?? []) as Express.Multer.File[];
    const saved: SavedFile[] = [];
    const errors: { name: string; error: string }[] = [];

    for (const file of files) {
      const result = await storeFile(grant, file);
      if ("error" in result) {
        errors.push({ name: file.originalname, error: result.error });
        continue;
      }
      saved.push(result);
    }

    res.status(saved.length === 0 ? 422 : 200).json({ ok: errors.length === 0, saved, errors });
  },
);

app.listen(3000);

Pay attention to the error messages. They're written for the model: on a 405 the agent sees the correct command, on a 401 it understands it just needs to request a new link. I covered this principle in more detail in "MCP: how to give AI access to your products and data".

Security: what the receiver should check

An upload link is a door into your server, and the agent is as untrusted a client as a form on your website. So:

  • Short lifetime. The link lives for 5 minutes. That's enough for the agent, and an old link from a log or a chat no longer opens anything.
  • The link is bound to one destination. A link for one record can't be used to upload a file to another. The client can't forge it or extend it.
  • An extension allowlist and a size limit. Accept only what your use case actually needs. Usually that's images and office documents.
  • Check the content, not the name. A text file renamed to .png is rejected with "This is not an image".
  • No SVG. An SVG only looks like an image and can contain a script. If files are served from a domain where users are signed in, that script gets access to their session.
  • The server builds the file name. The name from the request is used only to derive a key: transliterated, lowercase, hyphenated. A path like ../../.env goes nowhere.
  • Logging. Every upload is logged: which files went where.

Where it works and where it doesn't

The pattern works wherever the agent has a terminal and access to files: Claude Code, Cowork, Cursor, your own agent on a server. In a regular claude.ai or ChatGPT chat, the agent can't see the files on your computer and has nothing to send them with.

That's why I kept the old path. The MCP server instructions say: if you have the files and a terminal, upload them yourself; otherwise give the person a link for a manual upload. The agent picks the path on its own, and the workflow doesn't break in any client.

Honest limitations

  • The agent needs internet access from its terminal. In a sandbox with a closed network, the receiver's domain has to be allowed separately.
  • The client may ask for permission to run curl. That's expected: the command sends data out.
  • The link lives for 5 minutes. If the agent spent a long time looking for the files, it gets a 401 and requests a new one. That's an extra call, not an error.
  • A file with the same name replaces the old one. That's convenient for images in a text, but a document archive needs different logic.
  • No tokens are spent on the file's bytes, but the agent still doesn't see the image. If it needs to caption the image based on what's in it, it has to open the file separately.

Key takeaways

  • MCP tool arguments are written by the model. Everything that goes into them passes through the context and costs tokens.
  • Files don't belong there. MCP hands out a link or a rule, and the bytes travel over plain HTTP via curl.
  • Write the receiver's response for the model: ready-made lines to paste, a hint about what to do next, clear errors.
  • Keep a fallback path for clients without a terminal.

FAQ

Can you pass a file through MCP?

Technically, yes: you can encode the file in base64 and pass it as a string argument. But then the model pushes the whole file through its context, spends tokens and may corrupt the data. It's more reliable to have MCP hand out an upload link.

Why is this better than base64 in an argument?

The model writes one short command instead of hundreds of thousands of base64 characters. The file arrives byte for byte, and its size is limited only by your server.

Does this work in a regular Claude or ChatGPT chat?

No. In a chat, the agent has no terminal and no access to your files. There the agent gives the person a link, and the file is uploaded by hand in a browser.

Is it safe to give an agent an upload link?

No riskier than an upload form on a website, as long as the receiver is strict. The link lives for a few minutes, leads to one destination, and the server checks the type, size and content of every file.

Do I need a separate service for this?

No. The receiver is a single handler in the same backend where the MCP server lives.

If you want agents to work with files in your product without wasted tokens or manual uploads, get in touch. I'll help you design the MCP server and the file receiver.