Skip to content
Dylan Carter

The mahjong engine

A Hong Kong mahjong app: solo play against bots with no network, and cross-device multiplayer where each browser sees one seat and nothing else. This piece covers the move to an authoritative server and the claim-priority bug that every earlier test gate missed.

Play itRepo

Context

I was learning to play Hong Kong mahjong but when I searched for an app to play it on, they all sucked. From terrible UX, to english translations that made no sense at all, they were all flawed. So I decided to make a better one myself. I first started with making a working engine with some bots, then lessons, and then an AI coach.

The rules lived apart from the screen from the first commit. The engine is pure TypeScript with no React imports, the UI renders one player's view and dispatches actions, and the whole game state is plain data that survives a JSON round trip. Solo play runs offline in the browser. That separation cost little in week one and made the second phase possible without a rewrite.

One rule matters for this story. A discarded tile can be taken by other players: any seat may take it to complete three of a kind, called a pung, and the next seat in turn order may take it to complete a run of three, called a chow. A seat that can win off the tile outranks both. The rulebook is explicit about the order: win over pung, pung over chow.

Single player shipped first and still runs with no network at all. Multiplayer is the second phase.

The decision that mattered

The decision was to move authority to a server and keep one engine behind it.

The same room runner drives both modes through a transport interface. Solo hands it an in-memory transport and works on a plane. Multiplayer hands it a WebSocket transport into a Cloudflare Durable Object, one object per room. A Durable Object is single threaded, so the server's authority comes from the runtime's shape, with no locks to write or get wrong. The front end stayed on Vercel and the sockets live on Cloudflare; the browser bridges the two with one configured URL.

Authority means the browser is a dumb terminal. Clients send intents, and the server checks each one against the engine's own list of legal actions, then applies or rejects it, so a modified client gains nothing. Each browser receives its seat's view of the game and nothing more: the leak test records every payload sent to every seat across a full game and fails if one hidden tile appears. A parity test runs the same seeded game through the local transport and through a serialising network simulation and requires byte-identical output, which is the guard against the two modes drifting apart.

Claims are the distributed systems problem hiding inside the game. A claim window is the short pause after a discard in which the other seats may claim the tile. The server collects responses across the whole window and resolves them by the rulebook's ranking, with ties broken by seat order from the discarder. Arrival time carries no weight, so a fast connection cannot beat a slow one to a contested tile. Latency decides whether your claim arrives, and the rules decide who gets the tile.

The method was ten numbered gates, each with tests as its exit condition:

  1. Room runner and local transport with an injectable clock; solo rehosted on them
  2. Player-view audit and the leak test
  3. The Worker and the room object: lobby, seat tokens, a secret wall seed
  4. The authoritative loop over a real socket
  5. Claim windows and robbing the kong
  6. Turn timers, disconnect grace, bot takeover, reconnect
  7. Lobby and multiplayer table UI
  8. Coach policy, per-room rate limits, bring-your-own-key
  9. Integration suite and local-versus-network parity
  10. Multiplayer docs and deploy

What it cost

Gate five exposed a bug that phase one had shipped.

The gate's tests came from the rulebook's ranking table, and the pung-versus-chow case failed on an engine that had passed every gate before it. The resolver walked seats outward from the discarder and took the first seat holding any meld claim. A chow can be claimed by the nearest seat alone, so whenever a chow collided with a farther seat's pung, the chow won. The code had encoded nearest-first where the rule says rank-first.

The bug entered with the game state machine two days earlier. No test covered the collision, so every suite between stayed green, and any solo round played in that window resolved contested tiles by the wrong rule. The fix landed inside the gate with a regression test, mutation-checked so a faked fix kills the test.

I chose the other cost with the trade in view.

Turn clocks and claim windows run on an in-memory clock, as does the sixty-second grace before a dropped seat flips to a bot, and all of them re-arm from the persisted snapshot on restore. The one durable alarm belongs to room cleanup, because an abandoned room gets no traffic to wake it. The trade: a room evicted during a quiet stretch delays its deadlines until the next socket event. The server file names promoting gameplay deadlines to durable alarms as the known hardening step, and it remains open.

What I'd do differently

The engine split I would repeat without changes, and the parity test earned its place the first time the two transports disagreed in development. The change is where tests come from.

Every gate ended green because each gate's tests grew out of the code that gate added, and tests written from a diff inherit the diff's blind spots. The resolver never expressed the ranking, so no suite asked about it. The first tests derived from the rulebook's table found the bug in one case. I would write that adversarial suite before the reducer exists, while the rules are a table on paper and no code shape has had the chance to look plausible.

The timer deferral is the honest one. It was the right call for shipping, since the re-arm path covers reconnection, which is the failure phones produce all day. It also worries me more than anything else in the system, for the same reason the priority bug survived ten gates: the priority bug lived where no test looked, and the timer gap lives where my tests cannot look, because the fake clock in the suite cannot hibernate. A deadline that oversleeps an eviction will pass every test I can write today. That failure class deserves durable alarms sooner than "hardening step" suggests, and it is the first thing I would fund with another week.