Last week, in a moment of unbridled confidence that I can no longer explain, I offered to write a series on Effect Cluster. I don’t know why I did that, but I’m committed now I guess.
The problem is that cluster is far too large to cover in a single post, so this will be a series. That means that not every Module of the Week post will map cleanly to a single module.
I’m still calling it Module of the Week.
Deal with it.
Overview
This week we’ll stay at a higher level and focus on the architecture of cluster. Later posts will dig into how each piece works. To keep things concrete, we’ll ground our discussion in a hypothetical example app I’ve wanted for years. I apologize for it in advance.
The app
First, some context. I’m a dad. Like most dads, I have a condition. When someone sets up a dad joke, even by accident, I cannot let it go. I’ve tried. It comes out anyway.
At home, this is fine. My family has built up a tolerance. The problem is elsewhere. Slack, code review, GitHub issues, the neighborhood group chat. There’s a whole internet of people setting up perfectly good puns, and I just can’t be everywhere at once.
To address this problem, we will build DadBot - my sense of humor, packaged up as a bot you can bring into any chat.
-
Mike 10:41 AM
Anyone know why this test is flaky?
-
DadBot BOT 10:41 AM
I'd tell you a UDP joke, but you might not get it.
-
Mike 10:42 AM
...
-
Tim 2:17 PM
Can you split this into smaller commits?
-
DadBot APP 2:17 PM
I could, but I'm not ready to make that kind of commitment.
-
Tim 2:18 PM
I'm requesting changes ... to your personality.
DadBot makes one promise. He never tells the same joke twice in the same chat. No real dad could possibly keep that promise. Just ask my family.
But DadBot is going to live in millions of chats at once, with a dad in every one of them. Keeping that promise turns out to be a distributed systems problem, which is convenient, because that’s exactly what Effect Cluster was built for.
What you’ll learn
By the end of this post, you’ll know:
- What problem the actor model solves
- How Effect Cluster works as an actor system
- The pieces of a cluster: nodes, runners, shards, and entities
Shared state
Let’s start by building DadBot the obvious way.
Every chat has its own dad, and every dad remembers the jokes he’s told in that chat. When someone posts a message, the chat app calls a webhook, and our handler does three things:
- Loads the jokes he’s already told
- Asks the pun generator for one he hasn’t
- Saves the new list and post the reply
async function onMessage(chatId: string, message: string) { const told = await db.getJokesTold(chatId) const joke = await punGenerator.generate(message, { exclude: told, }) await db.setJokesTold(chatId, [...told, joke]) await chat.reply(chatId, joke)}Concurrent messages
This works if messages are handled sequentially. Our DadBot webhook handles them concurrently, so two handlers can read and write Dad’s joke list at the same time.
Click Play below to observe this behavior.
Slack #engineering
DadBot server
Dad · #engineeringActor
- Waiting for messages
- Mailbox
- –
- State
- Idle
Database Unlocked
| Joke told | Saved by |
|---|---|
| – | – |
| – | – |
| – | – |
Both handlers read DadBot’s joke list before either one saved, so both picked the same joke. We’ve already broken the sacred DadBot promise.
This is a race condition. Whether DadBot repeats itself depends on how the two handlers happen to line up, and the two seconds that it takes to generate a pun gives plenty of time for handlers to overlap.
Using locks
The usual way to stop two handlers from stepping on each other is a lock. A handler locks the chat’s joke list before reading it and unlocks it after saving. If another handler wants the list in the meantime, it has to wait.
Click Play below to observe this behavior.
Slack #engineering
DadBot server
Dad · #engineeringActor
- Waiting for messages
- Mailbox
- –
- State
- Idle
Database Unlocked
| Joke told | Saved by |
|---|---|
| – | – |
| – | – |
| – | – |
The lock fixes the repeats, but now DadBot replies three times, like someone who hasn’t read the rest of the thread. Clearly he needs to work on his delivery.
A lock can’t do that on its own. Each handler only knows about its own message. You would have to add more infrastructure, like a queue and maybe a scheduler that keeps track of busy chats.
And that’s all on one server. When DadBot goes viral, we’ll have to spread it across a fleet of machines, and messages from the same chat will land on different ones. The queue and the scheduler have to be shared between all of them. Every message waits on the database for a lock, and if a server dies mid-pun while holding one, the whole chat waits for it to time out.
That’s a lot of machinery just to know when to tell a joke … right?
What we want is for each chat to get its own Dad. One long-lived worker that receives every message, keeps its own memory, and decides for itself when to reply. Luckily, people have been solving distributed systems problems like this for over fifty years.
The actor model
In 1973, Carl Hewitt described the actor model. An actor is a small, independent worker with three properties:
- An address, so other code can send it messages
- A mailbox, where those messages wait their turn
- Some private state that nothing else can touch
An actor handles one message from its mailbox at a time, and the only way to change its state is to send it a message.
One actor per chat
Let’s give each chat its own Dad. His joke list becomes his private state, and every message in the chat will arrive in his mailbox. Internally, he will track the time between messages in the chat so that he can deliver a joke at the most opportune moment.
Click Play below to observe this behavior.
Slack #engineering
DadBot server
Dad · #engineeringActor
- Waiting for messages
- Mailbox
- –
- State
- Idle
Database Unlocked
| Joke told | Saved by |
|---|---|
| – | – |
| – | – |
| – | – |
By modeling the chat’s Dad as an actor with his own internal state, he can retain specific memory between messages and decide for himself when to talk.
Don’t let it go to your head.
Messaging by address
Actors also solve the scaling problem from earlier. When a message comes in from #engineering, which machine gets it?
The webhook doesn’t need to know. Every actor has an address, and that’s all you need to send it a message. The webhook sends the message to “Dad for #engineering”, and the actor system routes it to him, whichever machine he’s on. Actors message each other the same way.
Plenty of systems work this way. Erlang runs telephone switches on it, and Akka brought it to the JVM. Microsoft’s Orleans added virtual actors, which start the moment someone sends them a message. Nobody ever creates Dad for #engineering. You send him a message and he shows up, the same way a real dad appears the second anyone touches the thermostat.
That’s exactly how Effect Cluster works.
Effect Cluster
Effect Cluster is an actor system for Effect, but cluster calls its actors entities. They start the first time someone sends them a message, and passivate (shut down) once they’ve been idle for a while. More on that later.
DadBot has a single entity type, Dad, and every chat gets its own Dad entity. An entity’s address is its type plus an identifier, so “Dad for #engineering” is the entity with type Dad and identifier #engineering.
Architecture
So where does Dad actually live? A cluster starts with a few machines, called nodes.
Each node hosts one or more runners. A runner is your application: an ordinary Node.js or Bun process, started from your own code. It’s where your entity handlers actually execute. Every joke Dad tells gets dreamed up inside a runner.
To deliver a message, cluster needs to know which runner a Dad is on. Tracking that for millions of dads, one by one, would mean a huge routing table. So cluster groups entities into shards, a fixed set of buckets, and gives each shard to a runner. An entity’s identifier decides which shard it lands in, so to route a message, cluster only needs to know who owns each shard.
That means finding Dad comes down to two questions. Which shard is he in, and which runner owns that shard? We will cover that and more next time!
Wrapping up
That’s it for this week! Here’s what we covered:
The actor model: each chat gets its own Dad, with private memory and a mailbox he works through one message at a time
Effect Cluster: an actor system where actors are called entities, and you reach each one by its address, like “Dad for
#engineering”The pieces of a cluster: nodes host runners, runners own shards, and shards hold entities
Next week, we’ll dig into how cluster works out which shard Dad is in, and how the runners agree on who owns which shard without anyone in charge.
Until then, Happy Effecting!