Sandbox.

Interactive articles on proofs and data-infrastructure internals, then live simulations of the distributed-systems mechanisms I run in production. Every number on the page is computed, not typed.

Articles

Math and algorithms

Live simulations

One mechanism per card. Each is a real simulation — poke it: the point is to see the failure modes, not the happy path.

Streaming

01

Dataflow and backpressure

streaming · foundations

Three sources, a keyBy shuffle, two window operators, one sink. Hover an operator to slow it down: its queue fills, the edge into it drains slower, and the sources upstream of the shuffle back off. Click to pin.

Animated dataflow graph: three sources feed a keyBy shuffle and two window operators; slowing an operator fills its queue and throttles the sources.

02

Event time and watermarks

Flink · windows

Events arrive out of order. The watermark trails the maximum event time by a bound; a 5-second window fires when the watermark passes its end. Everything behind the watermark is late.

A timeline of out-of-order events with a trailing watermark; windows fire when the watermark passes their end and late events are dropped.

03

Checkpoint barriers

Flink · exactly-once

Two input streams, one operator. A barrier is injected into both; with aligned checkpoints the operator blocks the faster stream until the slower barrier arrives, then snapshots. Unaligned checkpoints skip the wait by snapshotting in-flight records.

Two streams into one operator with checkpoint barriers; aligned checkpoints block the faster stream, unaligned ones snapshot in-flight records.

Partitioning and rebalancing

04

Consistent hashing

foundations · ring · virtual nodes

240 keys, each node owns several ring positions. Add or remove a node and watch how few keys actually move — that's the whole point of the technique.

A hash ring with 240 keys coloured by owning node; adding or removing a node recolours only the keys between its positions.

05

Partitions and hot keys

Kafka · foundations

Keys are hashed onto 12 partitions, each drained at a fixed rate. Turn up the skew and one partition's lag runs away while the others idle — parallelism you can't use. Adding partitions re-hashes every key; about half land on a different partition, so per-key ordering does not survive it.

Twelve Kafka partitions as lag bars; a skewed key distribution piles lag onto one partition.

06

Consumer group rebalance

Kafka · foundations

Partitions are assigned to consumers. Add or remove one and count how many partitions change hands: the range assignor reshuffles almost everything, cooperative-sticky moves only what it must.

Twelve partitions assigned to consumers; a rebalance shows which partitions change owner under range versus cooperative-sticky assignment.

Consensus and replication

07

Raft

consensus · election · replication

Five nodes, one leader, election timeouts drawn as arcs. Kill the leader, or click any node to crash or restart it, and write entries: one commits once a majority stores it. Modelled: votes only for an up-to-date log, appends only onto a matching prefix, commits counted only for current-term entries. Not modelled: snapshots, membership changes.

Five Raft nodes exchanging heartbeats, votes and log entries; a killed leader triggers an election and entries commit on a majority.

08

Replication and quorum

Cassandra · consistency

Six nodes, replication factor 3: a key lives on three consecutive ring positions. Writes and reads succeed when enough replicas answer; a read returns the newest version among its replies. Kill a replica, write, revive it and read at ONE to see a stale read; R + W > RF rules it out.

Six replicas with replication factor 3; reads and writes succeed when the chosen consistency level of replicas answers.

W:R:

Storage

09

LSM tree

storage · memtable · L0 · L1 · L2

Writes land in a sorted memtable, flush to L0, and compact downward into non-overlapping tables. The number to watch is write amplification.

An LSM tree: a memtable flushing to L0 and compacting into non-overlapping L1 and L2 tables, with write amplification counted.

10

Bloom filter

probabilistic · foundations

256 bits, k hash functions. Inserts set bits; a query is "maybe" only if all its bits are set. No false negatives, ever — but watch the false-positive rate climb as the filter fills, and compare with the formula.

A 256-bit Bloom filter with k hash positions per item; queries show hits, misses and false positives against the formula.

11

Snapshots and time travel

Iceberg · table format

Every commit creates an immutable snapshot pointing at a set of data files. Old snapshots keep their files alive — until you expire them, at which point unreferenced files become garbage. Drag through history.

A chain of Iceberg snapshots over data files; expiring snapshots turns unreferenced files into garbage.

Batch execution and orchestration

12

Shuffle and stragglers

Spark · stages

Four map tasks write into three shuffle buckets; three reduce tasks pull from every map output. The wide dependency is the stage boundary — the whole job waits for the slowest map. Speculative execution launches a copy of the straggler.

Four map tasks, a shuffle boundary and three reduce tasks; a skewed map task holds the stage until speculative execution runs a copy.

13

Rolling update and rescheduling

Kubernetes · orchestration

Three nodes, a Deployment with N replicas. A rollout replaces pods one surge at a time and only after readiness passes. Kill a node and watch the scheduler place its pods elsewhere — or leave them Pending if nothing fits.

Three Kubernetes nodes running a Deployment; a rollout surges one pod at a time and a lost node sends its pods back to the scheduler.