Sorted Runs

LSM trees, from first principles to production systems

SlateDB: What if S3 Were Your Only Disk?

ep6slatedblsm trees3object storagefencing

The simplest database you can build on S3 gives every key its own object: one PUT per write, one GET per read. SlateDB’s launch post prices that design:

A naive system that uses S3 directly as a key-value store for a modest 10K ops/sec split even between reads and writes would cost $70K/mo and perform poorly.

Almost all of that bill is PUTs, which S3 charges at $0.005 per 1,000 against $0.0004 for a GET. SlateDB builds a real database on the same bucket without paying that, and keeps it safe with no coordinator, as long as the bucket really behaves like S3.

Some numbers come from “my test”, a small key-value service I run on SlateDB (one writer, four databases, compute on DigitalOcean in New York). It is not production. Its storage moved from DigitalOcean Spaces to Cloudflare R2 in June 2026, and every seven-day figure was measured after that. It runs a release older than v0.15.

A key-value library with a bucket for a disk

SlateDB is an embedded storage engine built as a log-structured merge-tree that writes its data to object storage. Keys and values are plain bytes kept sorted by key, and the interface is put, get, delete and a scan over a range. There is no server: it is a library linked into your program, like RocksDB, except that where RocksDB has a local disk, SlateDB has a bucket.

Chris Riccomini and Rohan Desai started it in March 2024 under Apache 2.0. The core is Rust, with bindings for Go, Java, Node and Python. Each database has exactly one writer process; any number of readers can open the same bucket from other machines without ever talking to the writer (RFC 0001 set that model). It also has transactions, merge operators, TTL on keys, and clones that start as a copy of another database without copying its files.

Underneath, it is the same tree as any other LSM: a memtable and a log, flushes to sorted files, background compaction. The design overview says “the only difference is that all of SlateDB’s data is written to object storage”. What changed is the address of every file: wal/, compacted/, manifest/, compactions/ and gc/ are folders in a bucket.

Three properties of the bucket drive every design choice (details): a request takes tens of milliseconds, every read and write is billed whatever its size, and an object is replaced whole or not at all.

Writes travel in batches

Writes collect in memory, and a timer ships them together as one numbered object in the write-ahead log, every 100 ms by default.

Many put calls flow into an in-memory WAL buffer holding 214 keys. A 100 ms timer ships the buffer as one numbered object, wal 39 dot sst, into the bucket, for a single PUT
Hundreds of writes become one numbered object per flush.

In one week my test wrote 358 million keys. They reached the bucket in 1.67 million log objects, about 214 keys per upload. Reads included, the request bill is about $45 a month at R2’s list prices. One object per key would be about $7,000 a month at the same prices (my arithmetic: 358 million PUTs a week at $4.50 per million, plus reads).

The price of a durable write

Batching is paid for in time. Of those 1.67 million uploads, three finished in 100 ms or less; 87% took between 100 and 250 ms. My test’s compute is in DigitalOcean and its storage in R2, so every upload crosses the public internet; the 50 to 100 ms figure assumes S3 in the same region.

So SlateDB can say OK at two moments. Once the write is in memory, the median in my test was 2.5 ms. Once it is in the bucket, a flush took 216 ms at the median and 885 ms at the 99th percentile. Since v0.16, put returns right away with a handle; waiting on it (await_durable) means the write is in the bucket. Not waiting is faster, but if the process dies before the next flush, that write is gone.

This is a triangle of latency, cost and durability, where you get two, and the flush interval is the knob:

Flush interval Worst-case wait Log PUTs per month, busy writer
100 ms 100 ms plus the upload about 26 million
1 s 1 s plus the upload about 2.6 million
60 s a minute of writes at risk on a crash about 43 thousand

The table is arithmetic from the default and its documentation, not measurement. My test sets the timer to 60 seconds and barely uses it. It reads from Kafka and flushes every database before committing its offsets, so a crash replays from the last commit and the durable log is upstream.

Fencing a writer with a file name

“Exactly one writer” is easy to say and hard to enforce, because a writer can freeze or lose its network, then come back still believing it is in charge. The bucket has no idea which process is the writer.

The list of files that make up the database lives in one object, the manifest: which sorted runs exist, where log replay starts, and two counters, the writer epoch and the compactor epoch. SlateDB never edits it. Each new version is a new object with the next number in its name.

Each version is created with If-None-Match: *: create this object only if it does not exist, which S3 has supported since August 2024. Two processes that both try to create manifest 8 cannot both succeed: one gets a 200, the other a 412.

Manifests 6 and 7 already exist. A writer and a compactor both try to create manifest 8 with If-None-Match star. The writer gets 200, the compactor gets 412, then re-reads manifest 8, applies its change and creates manifest 9
Losing the race is not an error: read the winner, apply your change on top, try the next number.

The loser re-reads manifest 8, applies its own change on top, and tries to create 9. That is how the writer and the compactor share one manifest without a lock.

The launch post calls the protocol based on conditional write IF-MATCH APIs, but the code creates manifests, log objects and compaction job files with the create-only header. If-Match appears only in the garbage collector’s boundary files, covered below.

The epoch and the zero-byte fence

Fencing is a separate step. A new writer that opens the database creates a manifest with the writer epoch plus one. The RFC defines a zombie as a writer whose epoch is less than the epoch in the current manifest. But how does the zombie find out?

The new writer also creates an empty, zero-byte object at the next number in the log. Log objects are numbered and created with the same header, so the zombie’s next flush tries to create that same object and gets a 412, which SlateDB maps to Fenced.

Writer B creates manifest 10 with writer epoch 5, then an empty wal 43 object. Writer A, still at epoch 4, flushes toward wal 43, receives 412, and closes as Fenced, with its pending durable writes returning errors
The fence is a file name: the zombie collides with it on its next flush.

The zombie closes itself as fenced, and every write still waiting to be durable gets an error instead of an OK that is not true. There is no lease and no coordinator: the zombie can work only until its next flush, and the new writer replays whatever it made durable before the fence.

It happened to me. In June, a Kubernetes rolling update started my new pod before stopping the old one. The two writers fenced each other in a loop, and with that day’s restarts the writer epoch climbed to 33. Nothing was corrupted, and the fix was one line: Recreate, which stops the old pod first.

Why the cache cannot serve a stale block

Every SlateDB process caches blocks in memory and, optionally, on local disk. Sorted-run files are named with a ULID, a unique ID that starts with a timestamp, and compaction writes its output under new names. A file can be deleted but never rewritten under the same name. The cache key includes the file’s ID, so a cached block is correct for as long as its file exists. Nothing needs invalidating.

Conditional writes fence writers, not caches. The one thing that does change is the manifest, and no process trusts its copy for long: by default the writer re-checks it every second and a separate reader every 10 seconds. A reader can be a little behind, never wrong.

Checking used to mean listing the whole manifest/ folder. RFC 0032 changed that in v0.15.0. Each process keeps the latest manifest bytes and asks for the next number instead: a GET for id N+1. A 404 means nothing changed, so it keeps its copy. After four consecutive hits it falls back to a LIST, and a cold store lists once. My test predates this release, so it still lists. The RFC’s motivation:

A database that’s idle for five minutes incurs 560 LIST requests with default configuration. This comes out to $24.19 per-month in AWS S3 us-east-1 pricing.

Left: the old way lists the whole manifest folder on every poll. Right: a process holding cached manifest 8 sends GET for manifest 9, gets a 404, and keeps its copy
Asking for the next number replaces listing the folder.

Bloom filters also keep reads off the bucket: in a week my test skipped 120 million file reads, with a 0.81% false-positive rate against the 10 bits per key default, and 98.8% of block lookups were served from memory.

Compaction as a job board

Compaction rewrites sorted runs into new ones, then changes the manifest to swap them in. So the compactor is a second writer of the manifest, with its own epoch checked the same way.

Since v0.14, compaction is split in two. A coordinator writes compaction jobs to a numbered file in the bucket; workers claim a job with the same create-only write. A worker sends a heartbeat every 10 seconds and holds no state; if it dies, its job returns to the board after 30 seconds, so workers can be cheap machines that might disappear. By default all of it runs inside the writer’s process, as in my test, where writes never stalled all week.

Deleting a name reopens it

Something still has to delete files: the garbage collector. But deleting a numbered file reopens its name. RFC 0026 describes the hazard:

A stalled writer can prepare file N+1, another writer can create and supersede N+1, GC can later delete N+1, and the stalled writer can then resume and successfully create the same filename.

So before deleting anything, the garbage collector raises a boundary, stored in a small file of its own, and deletes only files at or below it. After every successful create, a writer checks the boundary and treats a number at or below it as a failure.

The garbage collector raises gc slash manifest.boundary to 7, then deletes manifests 6 and 7. A stalled writer resumes and creates manifest 6, which succeeds because the name is free, but its check finds 6 is at or below the boundary and fails
The boundary file closes the gap that deleting a name opens.

The boundary files are the only objects SlateDB ever replaces, which makes them the only place that needs the other conditional header. If-Match means: replace this object only if its version tag, the ETag, is still the one I read. HTTP sends an ETag inside quotation marks, and S3 clients pass it back as received.

When the bucket is not quite S3

My test started on DigitalOcean Spaces in New York. A few weeks in, the bucket held 38 GB and the database inside it needed about 44 MB, so 99% of the bucket was garbage. Every write had succeeded and every read returned the right value. The garbage collector was running, retrying about 23 times a second, and had deleted nothing. Its boundary sat at 144 while manifests had reached 57,649.

The cause was one header, in the request that moves the boundary file. Create-only worked, so writes, reads and fencing were fine, but on June 15 the store returned 412 for every If-Match, even with the right tag. A week later the unquoted form started working, while the quoted one, the form HTTP specifies and real clients send, still failed:

PUT gc/manifest.boundary
If-Match: "a1b2c3"      ->  412 Precondition Failed
If-Match: a1b2c3        ->  200

I had checked this in March, before the test began, and it passed, partly because my test script stripped the quotes before sending the tag. The real client does not, so the check was testing a request nobody sends.

I did not wait for a fix: within days I moved storage to R2, which has supported conditional writes since 2022. Chris Riccomini opened an issue citing DigitalOcean, and within a few weeks a switch to turn the boundary files off shipped in v0.15. With the switch off, the only guard is the collector’s wait before deleting, so the documentation asks for a min_age longer than any stale process can live.

On July 9 DigitalOcean shipped a fix to one region: a new bucket in Richmond accepted the quoted tag, a new one in New York still refused it. Spaces is built on Ceph, and DigitalOcean runs 75 Ceph clusters for block and object storage, so my guess is that a fix can reach one region before another. On October 8, a bucket from June and one created in September, both in New York, still returned 412 for the quoted tag and 200 for the unquoted one.

It is not one provider’s mistake: a university research cloud on Ceph reported the same failure (one user’s measurement). And MinIO, the store many people test against locally, strips quotes before comparing, so laptop tests pass.

Loud failures are safe

Three rows. A store that rejects the header with an error fails loudly and the database refuses to open. A store that rejects only If-Match fails quietly, garbage grows, as in my test. A store that accepts and ignores the header lets two writers both win
The quieter the failure, the worse the damage.

A store that rejects the header with an error, as OVH did before it added support and Alibaba OSS documents, fails loudly: the database will not open. A store that rejects only If-Match fails slowly, like the one in my test. The worst case says OK and ignores the header: one user reports that Google Cloud Storage’s S3 interoperability layer does this. Then, by my reading of the create-only code, two writers can both create manifest 8 and both believe they won. The fence file is overwritten too, so the zombie never hears its 412, and acknowledged writes can disappear.

So before trusting a store that says it is S3-compatible, test both headers the way the real client sends them: create the same object twice (the second must get a 412), replace one with its quoted ETag (a 200), and replace one with a wrong ETag (a 412 again).

When to pick it, and when not to

Pick SlateDB when one writer is enough and you would rather pay per request than run disks, so any machine can die without losing a durable write. The launch post names Dropbox and ZeroFS among its users. Avoid it when a durable write must finish in a few milliseconds, when many processes must write to one database, or when readers must keep up with the writer: they see new writes only when they poll, every 10 seconds by default. The first limit has a way out: the log can move to a faster store, like S3 Express One Zone, and v0.16 opened the log interface to other backends (the RFC is still a draft; only the object-store log ships).

It is also still moving fast: v0.17.0 shipped on September 29, 2026, and the newest features carry the most risk. On October 2, PR #2132 merged a fix for a data-loss bug in union clones, which merge several databases into one:

Fix data loss after union clones combine L0 views with the same ID. Compaction can read one view but remove several.

The PR “only prevents the bug. It does not repair existing manifests or restore lost data.” Only databases built by union or projected clones are affected, and v0.17.0 still has the bug.

SlateDB has no disk, only a bucket and a few rules about file names, so every guarantee it makes is exactly as strong as the store’s implementation of two conditional headers.

Sources and further reading

Sources

  1. SlateDB, Introducing SlateDB, launch post, 2026-06-30.
  2. SlateDB, design overview and tuning guide.
  3. SlateDB, RFC 0001: Manifest, RFC 0002: Compaction, RFC 0025: Distributed compaction, RFC 0026: GC boundary, RFC 0030: Pluggable WAL, RFC 0032: Cached probing.
  4. SlateDB source: object_store.rs, writer_init.rs, wal store.rs, db_cache, config.rs.
  5. SlateDB releases: v0.16.0, v0.17.0; PR #1968, issue #1818, PR #1917, PR #2132.
  6. AWS, conditional writes and S3 pricing; Cloudflare, R2 pricing and release notes.
  7. IETF, RFC 9110 section 8.8.3, entity tags.
  8. DigitalOcean on Ceph: Ceph at DigitalOcean (2021) and Ceph Operations at Scale (2025).
  9. Provider reports: Ceph on Jetstream2, MinIO source, OVH, Alibaba OSS, GCS interop report.
  10. Telemetry labelled “my test” is from my own service, seven days ending 2026-09-27; the 38 GB, 44 MB, 23 retries a second, boundary 144 and 57,649 manifests figures are from my June incident notes.

Further reading

← all posts