S3: Millions of Hard Drives, and You Can't Edit One Byte
Corrections to the video
The article below has these right. The video does not.
- medium "$5,900 a month for a terabyte of RAM" is the full on-demand price of one r6i.32xlarge, CPUs included. The cheapest 1 TiB memory-optimized instance (x2gd.16xlarge) is about $3,900 a month, which makes the ratio to S3 about 166x, not 250x.
- low S3 storage is quoted as $24 per TB in one beat and $23 in another (binary vs decimal terabytes).
- low The "local disk" column mixes EBS gp3 prices (network storage) with local NVMe latency.
- medium For centroid-based vector search, the video says comparing the query to the cluster centers takes one round trip. The round trip downloads the centroids. The comparison happens locally.
In April 2026, AWS shipped S3 Files: mount an S3 bucket like a real disk. AWS says file-based applications run on your S3 data with no code changes, and the mount is plain NFS, no FUSE driver. That makes a wrong idea easy to believe: that object storage is a filesystem with a cheaper price tag.
It is a different primitive with different rules, and the rules decide what a database built on top of it looks like. Most database engines assume a local disk, a point the RUM conjecture makes visible. Swap that disk for S3 and storage becomes nearly free, but every request costs money.
You can seek on reads, never on writes
On a local disk you can seek to byte 4000 and overwrite four bytes. On S3 only the reading half of that works. A GET takes a Range header and downloads just that byte range. Writing has no equivalent. To change one byte you write the whole object again, as a brand-new object.
The PutObject reference puts the flip side in one sentence: Amazon S3 never adds partial objects. A write is the whole thing or nothing. The S3 Files launch post has the cleanest explanation: objects are like books in a library, and you can’t edit a page, you need to replace the whole book. The documentation of that same filesystem product states it flatly: S3 objects are immutable.
Two precision notes, because this is the claim most likely to draw a correction. A plain PUT to an existing key does overwrite it, as the conditional-writes page says. What you cannot do is edit bytes inside an object, only replace it. And this is about general purpose buckets: Express One Zone directory buckets accept appends to the tail of an object, but still no writes into the middle.
The irony is that S3 itself runs on millions of hard drives, and every one of them can seek. None of that seeking reaches your writes.
For more than a decade, engineers treated this as a problem to hide, and each of them paid for it. Others read it as a specification. Hold that thought, and keep a tab.
s3fs: the price was the bill
The first row on the tab is s3fs, from 2007. The promise was simple: mount the bucket, and every program treats it like a local disk.
People went to that trouble because of one number. A terabyte held in RAM costs about $5,900 a month: that is an r6i.32xlarge, a memory-optimized instance with 1 TiB of RAM, at $8.064 an hour on demand in US East (the whole machine, CPUs included). A terabyte on S3 costs about $24, $0.023 per GB, billed in binary gigabytes. That is roughly 250 times cheaper.
The number nobody quotes is the request count. A disk lists a folder for free. On a cold cache, s3fs needs a listing call per thousand objects plus a HEAD request for every object inside, to read the file metadata it stores in object headers. S3 charges per request: $0.0004 per 1,000 GETs and HEADs, $0.005 per 1,000 PUTs and LISTs. A million of those HEADs cost 40 cents. A billion cost $400. Every uncached directory listing becomes a billing event.
It leaked into the invoice. The project’s own FAQ has an entry titled Why does my AWS S3 bill cost a lot more than the storage fee? The answer points at updatedb walking the mount and making ListBucket and HeadObject calls.
That is the cost model in one sentence: with a disk you pay once and reads are free; with S3 storage is cheap and almost every operation is metered. The cost did not go away, it moved. You stop paying for bytes and start paying for verbs.
Goofys: the price was semantics
Goofys, from 2015, paid in a different currency. Its README jokes that it is a Filey System, not a File System, because it strives for performance first and POSIX second.
The README admits it cannot rename directories with more than 1,000 children but never says why. The source does. S3 has no rename on general purpose buckets (the RenameObject call exists only for Express One Zone), so Goofys copies every child one at a time, then deletes the originals in a single batch. That batch is DeleteObjects, which tops out at 1,000 keys.
This is not laziness. A rename that is atomic on a disk can die halfway through here and leave a partial copy next to the original. S3 Files has the same constraint and says what it does about it: S3 objects are immutable and do not support atomic renames, so on rename it writes the data to a new object and deletes the original. That is the move Goofys made in 2015. The primitive did not change.
A personal footnote: in 2015 I opened a ticket asking Ceph, the open-source storage system that speaks the S3 API, to make its multi-object delete faster. It deleted about ten objects a second on a 500-object request. A fix (concurrent deletes, reported at 5 to 6 times faster) was merged in December 2022. My ticket was marked resolved in September 2026. Two systems held the same fact and disagreed for almost four years.
JuiceFS: the price was a second database
JuiceFS, open-sourced in January 2021, got closest by pretending the least. It splits files into blocks (default maximum 4 MB), treats object storage as a dumb block store, and keeps the namespace in a separate metadata engine.
Listings got fast and renames got cheap. But a shared mount now runs two stateful systems instead of one. That is a second system to page you and a second one to back up. JuiceFS backs up its metadata to object storage every hour, but lose the database and those backups and the bucket is just numbered blobs: you cannot find the original file directly in the object storage.
What a replace-only primitive gives you
Three rows, three different prices, one cause: you can’t edit an object, only replace it. That sounds like a limitation. It is also a guarantee.
In August 2024 S3 turned that guarantee into a tool. Conditional writes create an object only if nobody has created it yet. Two nodes race for one key with If-None-Match: *. The first write to finish succeeds, and S3 fails the others with 412 Precondition Failed.
Whoever wins the race owns the key. Note the scope: this is the create-if-absent form. Compare-and-swap on an ETag with If-Match arrived separately, in November 2024.
Now put that next to the structure of an LSM tree, which buffers writes in memory and flushes them as sorted runs. Sorted runs are written once and never modified. Compaction does not edit them either: it writes a new run and drops the old ones. Object storage requires immutability, and the LSM tree already had it. Every filesystem wrapper had to solve how to edit a byte without replacing the whole object. An LSM never has to solve that problem at all.
The RUM triangle at S3 prices
The structure fits the primitive. The next question is what the primitive does to the trade-offs. Take the RUM triangle, which says an access method trades read, update and memory (space) overheads against each other (more on the RUM conjecture), and look at what changes.
Space stops mattering. On EBS, an extra terabyte of gp3 is $80 a month, and a replicated database pays it three times. On S3 it is about $23, with eleven nines of durability across at least three Availability Zones already in the price. Space amplification just stopped being the corner to worry about.
The whole stack tells the same story. turbopuffer publishes a cost table per terabyte per month: three SSD replicas with half the data cached in RAM run $1,600, S3 with an SSD cache runs $70. That is about 23 times cheaper.
Writes become money. On a local disk, write amplification wears out your SSD. On S3 every rewrite is a PUT, and a PUT costs $0.005 per 1,000 against $0.0004 for a GET, twelve and a half times as much. Write amplification turned into a bill, priced per object, not per byte.
The good news is that S3 offers exactly one kind of write, and it is the one an LSM makes. A flushed run, a compacted run and a finished log segment are all written whole.
Reads are the killer. On local NVMe, checking one more sorted run takes about a tenth of a millisecond. On S3, AWS documents first-byte latencies of roughly 100 to 200 milliseconds. On S3, read amplification is what kills you.
That is why the read half of the asymmetry matters. An LSM on S3 keeps its bloom filters and block index cached, then pulls only the block it needs with one ranged GET. Seek on reads is the only seek an LSM needs.
Where does the line fall? If your query can wait a second, reading straight from S3 works. turbopuffer’s own docs show a cold query at p50 874 ms against 14 ms warm, for a million documents. If a user is watching a spinner, you need a cache in front. Same data, different tolerance for waiting, different bill.
turbopuffer: designed around the bill
Simon Eskildsen spent eight years at Shopify, a very long time on call. In his CMU talk:
the only database that I would go on call for was one where I didn’t have to write the storage layer
So turbopuffer’s only stateful dependency is object storage. In that talk he says it serves more than three trillion vectors and documents in production, and that at launch it charged about a dollar per million vectors a month when incumbents charged close to a hundred.
The standard algorithm for vector search is HNSW, a graph you walk hop by hop. On S3 every hop is a GET, and each one costs about a hundred milliseconds. Say ten hops, and you have waited a full second (ten is an example, the hundred milliseconds per hop is from the talk).
turbopuffer never used that algorithm. Its vector index is based on SPFresh, a centroid-based index, which groups vectors into clusters. Downloading the cluster centers (centroids.bin) takes one round trip, and comparing the query against them happens in memory. Fetching the closest clusters, all at once, takes one more. Two round trips instead of ten. turbopuffer’s docs say the centroid design minimizes roundtrips and write amplification compared to graph indexes like HNSW.
Tonbo: let S3 pick the file format
turbopuffer let S3 pick its algorithm. Tonbo lets it pick the file format. It is an LSM tree on object storage: a WAL, a MemTable, and a flush that writes each run to S3 as a Parquet file.
The files your database writes are Parquet, a format DuckDB, Spark and Pandas already read. No export job, no second copy. Through Tonbo’s manifest there is one version of the truth.
The trade is on point lookups. A custom SSTable index lands on one small block of a few kilobytes. Parquet lands on a page built for scans, often megabytes, which makes single-row reads slower and scans cheap.
SlateDB: S3 as the only disk
Where Tonbo changed the format, SlateDB changes the write path. Chris Riccomini and Rohan Desai built it around one blunt idea: S3 as the only disk. The MemTable still lives in memory. Flushes and the write-ahead log go straight to S3. A local disk is at most a cache.
The write path. A durable write waits for its batch, then for S3 to confirm the PUT, which alone takes 50 to 100 milliseconds on S3 Standard. That is slow by local-disk standards. But once S3 says OK the write is durable, and every machine in your cluster can die while your data survives.
SlateDB also avoids paying a PUT for every write. Writes collect in a buffer, and every 100 milliseconds by default the whole batch goes out as a single PUT. That is its answer to the write corner: it spends fewer verbs.
Compaction, priced. Compaction is the other half of the bill. It reads runs with GETs and writes new ones with PUTs. Compact eagerly and you pay in PUTs. Compact lazily and every read pays in GETs instead. It is the RUM trade-off again, now with a price tag.
The conditional write comes back. If two SlateDB writers race for the next WAL entry, one wins and the other is fenced out: each WAL file is written exactly once, with the same create-if-absent primitive. No split-brain, no Raft cluster to run.
The read path. SlateDB answers the read corner with caching: memory holds the hottest blocks, an optional local disk cache holds warm ones, and S3 holds the rest. Hot reads never touch S3 at all.
That settles JuiceFS’s row on the tab. JuiceFS put its metadata in a second stateful system. SlateDB keeps its metadata, the manifest, on S3 itself. One system fewer, not one more.
One exception: for a faster log, SlateDB can put its WAL on a separate object store such as S3 Express One Zone, at single-digit milliseconds. PUTs there cost about 77 percent less and storage costs about five times more per GB. Cheap verbs and expensive bytes is exactly what a write-ahead log wants.
It has a name: diskless
Building with S3 as the only disk now has a name. SlateDB’s June 2026 introductory post says “diskless” systems that delegate durability to object storage are the future of database systems. The pattern is wider than one project: ClickHouse Cloud’s compute nodes dropped their local disks, InfluxDB 3 can run diskless with object storage as the only persistence layer, and WarpStream built Kafka on object storage, then was acquired by Confluent.
Two filesystems on S3, opposite architectures
Back to the tab. In July 2025 Pierre Barre released ZeroFS: a bucket served as a POSIX filesystem over NFS, built on SlateDB. Nine months later AWS shipped S3 Files, which also serves files over NFS.
Both solve the same problem with opposite architectures. AWS built theirs on EFS and kept one file as one object, so the bucket still reads normally. ZeroFS gave that up: file contents are split into extents and packed into its own compressed segment objects, with metadata in an LSM tree. Since its 2.0 storage engine, the LSM holds metadata only.
Close the tab. s3fs paid in dollars. Goofys paid in semantics. JuiceFS paid with a second database. ZeroFS paid with a bucket you cannot read as files, but underneath, it stopped pretending: it calls itself a log-structured filesystem for S3.
And AWS’s answer? Andy Warfield of the S3 team wrote that rather than trying to hide it, the boundary itself was the feature we needed to build. To give you a filesystem, they put a real one in front of the bucket. The databases built for S3 never tried to hide the boundary at all.
Same map, new prices
The LSM recipe still holds (buffer, flush, compact), but the disk is now S3, a flush is a PUT that costs money, and compacting means deciding which PUTs are worth paying for. Put the three systems back on the triangle. turbopuffer lives by the read corner. SlateDB lives by the write corner. Tonbo sits between read and space. The trade-offs still hold and you still cannot have it all. Only the prices changed.
Long before any of these existed, one engine was already running LSM trees at Facebook scale on local disk, built around exactly the read, write and space amplification trade-offs we have been pricing here. That engine is RocksDB, covered in RocksDB, the SQLite of storage engines. The ideas it perfected on local disk are the ones these systems are now repricing for S3.
Sources and further reading
Sources
- AWS, S3 pricing
- AWS, EC2 on-demand pricing
- AWS, EBS pricing
- AWS, Amazon S3 now supports conditional writes
- AWS, How to prevent object overwrites with conditional writes
- AWS, Optimizing Amazon S3 performance
- AWS, S3 Express One Zone
- AWS, S3 storage classes
- AWS, PutObject API reference
- AWS News Blog, Launching S3 Files
- AWS, S3 Files synchronization
- AWS, Mounting S3 Files
- Andy Warfield, S3 Files and the changing face of S3, 2026
- Sirupsen, napkin-math
- s3fs-fuse, FAQ
- kahing, Goofys
- JuiceFS, Architecture
- Pierre Barre, ZeroFS
- Ceph, Feature #11326: rgw: make MultiDelete faster, PR 48679, RADOS Gateway docs
- Athanassoulis et al., Designing Access Methods: The RUM Conjecture, EDBT 2016
- Simon Eskildsen, CMU database talk on turbopuffer
- turbopuffer, launch post and cost table and architecture
- Malkov and Yashunin, HNSW
- Xu et al., SPFresh, SOSP 2023
- Tonbo, tonbo-io/tonbo
- SlateDB, slatedb.io and slatedb/slatedb
- Riccomini and Desai, Building a cloud-native LSM on object storage
- ClickHouse, ClickHouse Cloud stateless compute
- InfluxData, InfluxDB 3 OSS GA
- WarpStream, Kafka is dead, long live Kafka
- Facebook Engineering, Under the Hood: Building and open-sourcing RocksDB, 2013
- Dong et al., RocksDB: Evolution of Development Priorities, ACM TOS 2021
Further reading
- Andy Warfield, Building and operating a pretty big storage system called S3, 2023
- AWS, GetObject, DeleteObjects and RenameObject API references
- AWS, conditional compare-and-swap with If-Match, November 2024
- JuiceFS, open-sourcing announcement and metadata backup
- Goofys, directory rename source
- SlateDB, An Object-Native LSM for Online Systems (2026), tuning, manifest and fencing RFC and separate WAL store RFC
- Tonbo, Introducing Tonbo