LibraryGit Articles10 min read
SnapOtter and the file you upload to a stranger
The 240 MB walkthrough video that will not attach to an email and the scanned lease nobody can search get handled the same way in most small companies: whatever free converter site came up first. SnapOtter runs 200 of those tools on one container inside your own network, and the honest cost is not the subscription you cancel.

The 240 MB walkthrough video that will not attach to an email, the ninety-page scanned lease nobody can search, and the eleven supplier PDFs somebody has to merge before the Friday billing run all get handled the same way in most small companies. Whatever free website came up first, or a small subscription nobody remembers approving. CloudConvert meters that work by the minute, starting around eight dollars a month for a thousand conversion minutes, with one-time credit packages from about nine dollars for five hundred. Smallpdf lists Pro at fifteen a month billed monthly, ten if you commit for the year. The free sites charge nothing, which is the part that should worry you, because your signed lease went to one of them last quarter and no one can tell you which.
SnapOtter does that entire category of work on a server you control, and the install is one docker run.
What the repo actually is
SnapOtter is an AGPLv3 file-processing stack sitting at roughly 2,700 stars, with about 2,289 commits behind it and new ones landing today. That last part matters more than the star count. The 2.0 release in July rebuilt the thing on Postgres 17 and a Redis-backed job queue specifically so that long jobs, video transcodes, image upscales, OCR passes over a fat scanned document, survive a browser tab closing. Version 2.2.0 followed at the end of July. Commits have kept coming through this week, including dependency bumps for image-library advisories and a fix to how the transparency tool reads an alpha mask.
The catalogue is 200 or so tools across five kinds of file. Images get the biggest share at 107, covering resize, crop, compress, convert, watermark, vectorize, duplicate finding, passport photos, and the specific format conversions people actually search for, HEIC to JPG being the one that eats an hour of somebody's week every time a phone photo lands in a shared drive. Video gets 57, including convert, compress, trim, stabilize, burn and extract subtitles, and pull the audio out. Audio gets 27. PDF gets 29: merge, split, compress, redact, sign, watermark, page numbers, OCR. Then there are 23 file tools that are quietly the most useful ones in an ops context, because they do CSV to JSON, Excel to CSV, CSV merge and split, and ZIP handling.
On top of that sits local AI that runs on your hardware rather than somebody's API: background removal, image upscaling, photo restoration, object erasing, face blurring, text extraction, audio transcription, and auto-generated video subtitles. There is a layer-based image editor in the browser. There is OIDC login, so it plugs into Google or Okta or whatever you already use rather than becoming another password. There is a REST API with key auth and interactive docs, and a pipeline builder that chains tools into a reusable workflow up to twenty steps deep and exports as JSON.
The quick start is genuinely one line, a single container with Postgres and Redis embedded, and you log in at admin / admin and get told to change it. The production path is a three-container Compose file the README prints in full. It runs on AMD64 and ARM64, which means an old Intel box, an Apple Silicon Mac, or a Raspberry Pi. The license file is the real AGPLv3, not a source-available license wearing its clothes, and the project is dual-licensed with a paid commercial option for anyone who wants to put it inside a product they sell.
Why this lands differently than a cheaper subscription
For the owner-operator, the money here is not one invoice you can cancel. That is exactly why it never gets addressed. It is fifteen dollars on somebody's card for the PDF tool, a CloudConvert credit package from the year somebody automated a batch of file conversions, and a long tail of free sites that cost nothing and appear on no report. Add it up across a 40-person company and you get a number worth maybe two or three hundred a month, which is not a crisis. Fine. The reason to care is not the two hundred.
The reason to care is that the free site is where the expensive thing happens. Your CRM has a data processing agreement. Your accounting platform has one. The website that compressed the executed lease so it would fit under the email attachment limit has terms of service that your operations lead did not read, no relationship with your business, and no obligation to tell you anything. Nobody decided this. It happened because the file was too big and the deadline was that afternoon, and the honest truth is that it will keep happening until the internal option is faster than opening a new tab.
For the RevOps or marketing-ops reader, the API and the pipelines are the actual story, and they replace a different currency. The recurring chore of resizing and renaming every product image, or pulling the audio off thirty field recordings, or converting a supplier's CSV into the JSON shape your automation expects, is precisely the kind of request that gets filed as a ticket, ranked below the thing engineering is actually working on, and sits there for six weeks. A tool catalogue with an HTTP endpoint per tool and an API key is something an ops person can wire into the automation platform they already run, on a Friday, without asking anyone. That is the same displacement the operator gets, just paid in queue time instead of dollars.
The honest take
Start with the hours, because "one docker run" is true and also misleading. You will have a working instance on a laptop in about twenty minutes, and it will feel finished. It is not finished. A deployment you would actually put a company's files through means the Compose stack rather than the single container, a real Postgres password instead of the one printed in the README, TLS in front of it, the TRUST_PROXY setting matched to your actual network, a backup of the Postgres volume that you have restored at least once, and a decision about whether this is reachable from outside the office at all. That is a half day if you have deployed something like it before, and closer to a full day plus a second session if you have not. Budget the second session. Everyone skips it and everyone regrets it.
Then the hosting, which is not the five dollar box. Transcoding video and running OCR over a long scanned document are both genuinely CPU-hungry, and the queue exists because these jobs take minutes. Four vCPUs and 8 GB of RAM is the honest floor for a small team, which runs about sixteen or seventeen euros a month at Hetzner after this year's price increases and forty-eight dollars a month at DigitalOcean for the equivalent shared-CPU droplet. Add storage, because file processing generates files. Call it twenty to sixty a month depending on where you host and how much you keep, against the two hundred or so in scattered subscriptions. The savings are real and they are not dramatic. The privacy change is the dramatic part.
What the paid products still do that this does not: they work from a phone, on hotel wifi, with no VPN, for a person who is never going to learn a new tool. That is not a small thing. Your self-hosted instance lives inside your network, and unless you deliberately expose it, the tech standing on a roof with a photo to compress cannot reach it. Solving that means SSO and a public hostname, which is more work and more exposure, and you should decide that on purpose rather than discovering it three weeks in. CloudConvert also supports a wider spread of obscure formats and will absorb a burst of ten thousand conversions without you thinking about it. And to be clear about scope, transcribing an audio file you already have is not the same job as a meeting notetaker that joins the call and writes it down. This does the first one.
Now the part most write-ups of a repo like this will skip. The 2.2.0 release notes are mostly security. Five advisories in one release, including a privilege escalation where a non-admin holding a delegated users:manage permission could reset a built-in administrator's password, a path traversal through malicious 1.x imports that could read or delete files outside the storage directory, an authorization gap that let users cancel other users' jobs, and a rate-limit bypass caused by the Docker default trusting every peer's forwarded headers. Nine more CVEs were patched in dependencies. The maintainers' own explanation of the endpoint authorization bug is the most useful sentence in the whole document: "All 45 hand-written routes had to remember the same call and none of them did."
Read that correctly. It is not a reason to walk away, and a project that publishes numbered advisories with real write-ups is behaving better than most commercial vendors in this category, who patch the same class of bug quietly. It is a statement of what you are signing up for. The moment you move the file errands inside your network, you own an authenticated multi-user service with an upload endpoint, an API, and a job queue, which is a meaningfully larger security surface than a bookmark to a website. The real failure mode is not the bug they already fixed. It is the instance somebody stands up in March, forgets about, and never updates, still running the version with the privilege escalation in it while the company keeps feeding it contracts.
Who maintains it matters too. The README says plainly that SnapOtter is built and maintained independently, with no venture capital or corporate backing, funded by sponsorships, and contributions require a CLA. The dual license tells you the rest of the business model: the commercial license is how this is supposed to pay for itself. The practical read is that if that revenue does not show up, the release pace slows before anything is ever announced. Your container will keep running either way, because that is the nice thing about software you host. The patches are what stop. So treat the exit as part of the plan: your files are on your own disk in ordinary formats, which means leaving costs you a weekend rather than a migration project, and that is the property worth protecting when you choose something like this.
One more thing worth knowing before you decide, because it is the kind of detail a vendor page buries. Files never leave your network, and that claim holds up in the architecture. Basic usage analytics, however, are on by default. The project says so openly and gives you two ways off, a build-time SNAPOTTER_ANALYTICS=off and an in-app admin opt-out. Nothing about your documents is going anywhere. Telemetry about which tools get used is, until you turn it off. If your written policy says no outbound telemetry, that is a five minute task on day one, not a discovery in an audit.
The AGPL question resolves more simply than people fear. Run it internally, modify it all you want, and you owe nobody anything. Run a modified version as a network service that other people use and you have to publish your changes. Put it inside a product you sell and you need the commercial license. For an ops team standing this up behind their own login, the license is a non-event.
The subscription was never the expensive part. The expensive part was that for years, the fastest way to handle your own files was to hand them to someone who had no reason to tell you what happened next.
Sources
Every claim above traces back to one of these. Go read them yourself.
- 01SnapOtter, open-source self-hosted file-processing infrastructure
snapotter-hq / github.com / retrieved Sep 17, 2026
- 02SnapOtter v2.2.0 release notes
snapotter-hq / github.com / retrieved Sep 17, 2026
- 03CloudConvert pricing
CloudConvert / cloudconvert.com
- 04Smallpdf pricing
Smallpdf / smallpdf.com
- 05DigitalOcean Droplet pricing
DigitalOcean / digitalocean.com
- 06Hetzner Cloud pricing
Hetzner / hetzner.com
Suggested reading
Selected articles based on topic, tags, and skill focus across the library.
Git Articles
OpenMAIC and the authoring seat you use twice a year
Turning the safety manual into actual training costs $1,749 a year for one Articulate seat that somebody opens twice, or five figures if you hand it to a course shop. OpenMAIC takes the PDF and builds the course, MIT licensed, and it shipped a release yesterday.
Git Articles
Pascal and the floor plan you pay for before you sign
A drafted floor plan runs $500 to $2,000, and the do-it-yourself route is a SketchUp Pro seat at $399 a year. Pascal is an MIT-licensed 3D building editor that opens in a browser tab with one command and takes instructions from the coding agent you already pay for. It also has three contributors and a stable release from June.
Git Articles
AutoClip and the minute you pay for either way
A clipping subscription charges one credit per minute of footage you upload, whether it finds anything worth keeping or not. AutoClip runs the same pipeline on the laptop you already own, MIT licensed, and as of Sunday it installs by double-clicking a file.

