Document Packets: Many Files, One Document, Without the Glue Code
A contract arrives with two annexes. An invoice arrives with the delivery note that justifies it. The document is the bundle, not any one file. A document packet is how anyformat runs a workflow on that bundle: one upload, one run, one set of extracted fields, with the context of every file in view.
Document processing tools, ours included for a long time, share one quiet assumption: a document is a file. One PDF in, one result out. It holds for most of what a back office handles, which is why the assumption survived as long as it did.
Then you look at the cases where it does not hold, and they turn out to be the expensive ones. A supply contract whose price schedule lives in Annex II, sent as a separate PDF. An insurance claim that is a form, three photos and a repair quote. A supplier invoice that only makes sense next to the purchase order and the delivery note. A loan application that is the application itself, two payslips and a bank statement. Nobody who works these files thinks of them as four documents. They are one case that happens to be spread across four files.
Why file-by-file processing breaks on these
Run each file on its own and every file comes back with a partial answer. The contract has the parties and the term but not the prices. The annex has the prices but no idea which contract it belongs to. The extraction for each file is correct and the answer to the actual question is nowhere.
So teams write the glue. Upload each file separately, keep a table that says which file belongs to which case, run them one by one, then merge the results in their own code: this field from the contract, that one from the annex, and a rule for what happens when both files mention the same date. It works until a case arrives with the files in a different order, or with a third annex, or with the annex before the contract. The merge logic is the part of the pipeline that nobody tests and everybody depends on.
The other workaround is to concatenate the PDFs into one before uploading. It keeps the context together and throws away everything else: the file names, the boundary between documents, and the ability to tell anyone which original file a value came from.
Document packets: the bundle is the unit
A document packet is the unit an anyformat workflow runs on: one or more files that the platform treats as a single document. Most of the time a packet holds one file, and you never notice it exists. It becomes visible the moment several files belong together.
Extraction is scoped to the packet, not to the files inside it. Parse and Extract see the whole bundle at once, so the price in the annex and the party in the contract are read in the same context. The result comes back at the packet level: one set of extracted fields for the whole document, with confidence and evidence on each field, the same shape a single-file run returns. Classify, Split and Validate take the packet as one input too, so a validation rule can compare a total on the invoice with a quantity on the delivery note without any code of yours in between.
The files keep their identity inside the packet. Each one has its own id and its own name, so when you need to look at one file on its own, you can.
One call from bytes to a result
Creating a packet is an upload. A single request takes between 1 and 100 files and turns them into one packet, atomically: either every file is stored or nothing is. There are two ways to hand over the bytes. A multipart upload, when the files sit in your application. Or a list of HTTPS URLs, such as presigned links to your own object storage, which anyformat fetches server-side so nothing streams through your backend. The URL import is all or nothing as well: one failed fetch and no partial packet is left behind.
Upload and run can be one call. Retries are safe: send the same Idempotency-Key and you get the original packet and the original run back, with no duplicate upload and no second extraction billed. And a packet can be run again at any point on the latest version of its workflow, without uploading anything. Every run is kept, so the results from before a schema change are still there to compare against.
The context you already have
Most documents arrive with facts your own systems already know. The customer the contract was signed for. The batch it came in. The vendor id you fetched it for. Packets carry that context with them: every upload accepts an optional metadata object, free-form JSON, stored with the packet and echoed back exactly as you sent it.
Metadata is not only a label. When the packet runs, Extract sees it. If a top-level key has the same name as a field in your schema, the value you passed is used for that field in preference to anything read off the page, and the field's evidence says so, with the marker metadata.<key> instead of a quote from the page. Both kinds of value sit in the same result, and your code can always tell which one your system supplied and which one was read from the document. A metadata value is also kept out of the confidence scoring, so it cannot inflate a score it did not earn.
Because metadata enters the prompt, it is bounded before it gets there. Control characters are collapsed, each value is capped at 512 characters and the whole block at 8,192. The copy you read back from the packet is still the original.
When to bundle, and when not to
The rule is simple. Bundle files into one packet when they are one document for the purpose of extraction: a contract and its annexes, a claim and its evidence, an invoice and its supporting paperwork. Upload unrelated files as separate packets, which keeps their results independent. A packet is a statement that the files belong together, and the result treats them that way.
There is one current limit worth knowing: a smart-table field cannot run over a packet that holds more than one file, and the run result says so explicitly instead of returning an empty table.
Nothing about billing changes. Usage is counted per page, and the pages of a packet are the pages of its files. A 12-page contract with two 3-page annexes is 18 pages, so a Parse and Extract workflow over it costs 60 credits per page, 1,080 credits for the whole packet, exactly what the three files would cost on their own. The difference is that you get one answer instead of three halves.
Every run in anyformat is a run on a document packet today, and multi-file packets are created through the v3 upload endpoints. If you integrated against v2, the object you know as a file collection is the same thing under a new name, and the ids carry over. The document packets page covers the model in full, and the upload-and-run endpoint is the shortest path from a folder of files to one set of fields.
anyformat is the document intelligence platform that turns unstructured documents into reliable, structured data, with enterprise-grade security, confidence scoring, and full auditability. Learn more at anyformat.ai.

