Rust Typestate: Designing APIs That Make Invalid States Unrepresentable
Learn how Rust typestate uses the type system to model object lifecycles and prevent invalid operations at compile time, with practical PhantomData examples.
Rust Typestate: Designing APIs That Make Invalid States Unrepresentable
Most programming languages let you create an object first and worry about whether it is actually ready to use later.
A connection might be disconnected.
A file might be closed.
A transaction might already be committed.
A parser might not have received enough input.
A device might not have been initialized.
In many systems, these conditions are represented by booleans, nullable fields, runtime checks, or conventions documented somewhere that developers may forget.
Rust gives us another approach:
make the state part of the type itself.
This idea is commonly called the typestate pattern.
Instead of creating one type that can exist in many vaguely valid states, we create different compile-time states and allow only the operations that make sense for each state.
The result is an API where the compiler becomes part of the protocol design.
The Core Idea
Imagine a connection with two states:
Disconnected → Connected → Disconnected
A traditional implementation might look like this:
struct Connection {
address: String,
connected: bool,
}
impl Connection {
fn send(&self, message: &str) {
if !self.connected {
panic!("connection is not active");
}
println!("Sending: {message}");
}
}
The API allows send() to exist even when the connection is disconnected.
The programmer has to remember the rule:
“Only call
send()after connecting.”
Rust can express that rule directly in the type system.
Turning State Into a Type
We can represent each state with an empty type:
struct Disconnected;
struct Connected;
These types do not need to store any runtime information.
They simply represent two different compile-time states.
Now we can make the connection generic:
use std::marker::PhantomData;
struct Connection<State> {
address: String,
_state: PhantomData<State>,
}
The State parameter tells Rust which state the connection belongs to.
PhantomData<State> is useful here because it lets the type carry that relationship even though the state itself does not need to occupy storage. Rust documents PhantomData as a zero-sized type used to communicate type-system relationships such as ownership, variance, and related compiler analysis.
Conceptually:
Connection<Disconnected>
Connection<Connected>
are two different types.
That difference is extremely important.
Restricting Methods by State
Now we can implement methods only for the states where they make sense.
impl Connection<Disconnected> {
fn new(address: impl Into<String>) -> Self {
Self {
address: address.into(),
_state: PhantomData,
}
}
fn connect(self) -> Connection<Connected> {
Connection {
address: self.address,
_state: PhantomData,
}
}
}
connect() exists only for:
Connection<Disconnected>
It consumes the disconnected connection and produces:
Connection<Connected>
Then we can define operations for the connected state:
impl Connection<Connected> {
fn send(&self, message: &str) {
println!("Sending '{message}' to {}", self.address);
}
fn disconnect(self) -> Connection<Disconnected> {
Connection {
address: self.address,
_state: PhantomData,
}
}
}
Now send() is available only when the compiler knows that the connection is connected.
Using the API
The normal flow looks natural:
fn main() {
let connection = Connection::<Disconnected>::new("server.example");
let connection = connection.connect();
connection.send("Hello Rust!");
let connection = connection.disconnect();
}
The important part is not that this code works.
The important part is what cannot be written successfully.
Consider:
fn main() {
let connection = Connection::<Disconnected>::new("server.example");
connection.send("Hello!");
}
This fails because the method belongs to:
impl Connection<Connected>
not:
impl Connection<Disconnected>
The compiler rejects the invalid state transition before the program runs.
The Compiler Becomes a Protocol Checker
This is where typestate becomes particularly interesting.
A normal API often documents a protocol like this:
create
↓
initialize
↓
start
↓
use
↓
stop
But documentation alone does not enforce the sequence.
Typestate can encode the sequence directly:
Device<Created>
│
▼
Device<Initialized>
│
▼
Device<Running>
│
▼
Device<Stopped>
Each state can expose a different set of operations.
For example:
struct Created;
struct Initialized;
struct Running;
struct Stopped;
Then:
impl Device<Created> {
fn initialize(self) -> Device<Initialized> {
// ...
}
}
and:
impl Device<Initialized> {
fn start(self) -> Device<Running> {
// ...
}
}
and:
impl Device<Running> {
fn stop(self) -> Device<Stopped> {
// ...
}
}
The resulting API resembles a state machine, but the compiler enforces the legal transitions.
Why self Is Often Consumed
You may notice something important in these methods:
fn connect(self) -> Connection<Connected>
rather than:
fn connect(&mut self)
Consuming self is powerful because the old state disappears.
Consider:
let connection = Connection::<Disconnected>::new("server");
let connection = connection.connect();
The original value of type:
Connection<Disconnected>
has been moved.
The new value has type:
Connection<Connected>
We have effectively transformed one state into another.
This fits naturally with Rust's ownership model.
Instead of:
same object
↓
boolean changes
we get:
old type
↓
ownership transfer
↓
new type
The state transition becomes a type transition.
The Interesting Part: Zero Runtime State
One of the most useful properties of this technique is that the state marker does not have to become a runtime field.
For example:
struct Connected;
struct Disconnected;
contain no data.
And:
PhantomData<State>
is zero-sized.
So the type:
Connection<Connected>
can carry compile-time information without requiring a boolean such as:
connected: bool
The state exists primarily for the compiler.
That gives us an interesting design principle:
Some information belongs in the type system, not in runtime memory.
Typestate vs Runtime Validation
Typestate does not eliminate runtime validation.
That distinction is important.
Suppose connecting to a server can fail because the server is unreachable.
The compiler cannot know whether the network connection will succeed.
So the transition can still return a Result:
impl Connection<Disconnected> {
fn connect(self) -> Result<Connection<Connected>, ConnectError> {
// Perform runtime connection attempt.
Ok(Connection {
address: self.address,
_state: PhantomData,
})
}
}
Now we have two layers:
Compile time
↓
Is this operation legal for this state?
Runtime
↓
Did the operation actually succeed?
These are different problems.
Rust can handle both.
A More Realistic Example: Transactions
Transactions are a great example.
Imagine these conceptual states:
Transaction<Active>
Transaction<Committed>
Transaction<RolledBack>
You might want:
impl Transaction<Active> {
fn commit(self) -> Transaction<Committed> {
// ...
}
fn rollback(self) -> Transaction<RolledBack> {
// ...
}
}
Meanwhile:
impl Transaction<Committed> {
// No commit() method.
}
Now code cannot accidentally commit the same transaction twice through the same typed API.
Likewise, a rolled-back transaction does not expose operations intended for an active transaction.
The state machine becomes visible in the API.
Typestate Is More Than PhantomData
It is tempting to think:
“Typestate means using
PhantomData.”
Not exactly.
PhantomData is one tool for representing compile-time relationships.
The deeper concept is:
Use types to encode valid states and legal transitions.
You can often design typestate APIs without PhantomData, especially when state is naturally represented by distinct structs.
For example:
struct Locked {
key: String,
}
struct Unlocked {
key: String,
}
You could transition between completely different structs:
impl Locked {
fn unlock(self, password: &str) -> Option<Unlocked> {
if password == "correct" {
Some(Unlocked { key: self.key })
} else {
None
}
}
}
Here the state is represented by the actual Rust type itself.
PhantomData becomes especially useful when several states share the same underlying representation and only the type-level state changes. Rust's documentation also distinguishes phantom type parameters from PhantomData: a phantom type parameter is unused at runtime, while PhantomData provides a compiler-visible connection to a type.
Designing State Machines With Rust Types
A useful way to approach typestate design is to first draw the state machine.
For example:
┌──────────────┐
│ Disconnected │
└──────┬───────┘
│ connect()
▼
┌──────────────┐
│ Connected │
└──────┬───────┘
│ disconnect()
▼
┌──────────────┐
│ Disconnected │
└──────────────┘
Then ask:
Which operations are valid in each state?
For our connection:
Disconnected
└── connect()
Connected
├── send()
└── disconnect()
Finally, encode those rules into impl blocks.
This turns API design into state-machine design.
A Useful Mental Model
Think of each type as a permission token.
Connection<Disconnected>
means:
“This value has the permissions and operations associated with the disconnected state.”
While:
Connection<Connected>
means:
“This value has the permissions and operations associated with the connected state.”
Changing the type changes the permissions.
That makes Rust's type system feel less like a collection of syntax rules and more like an API security boundary.
Where This Pattern Shines
Typestate becomes especially valuable when incorrect ordering can produce subtle bugs.
Examples include:
Database transactions
File lifecycle APIs
Network protocols
Hardware drivers
Resource initialization
Cryptographic protocols
Parsers
Connection pools
Builder APIs
Embedded systems
Concurrency abstractions
A builder is another familiar example.
Instead of allowing:
builder.build()
at every stage, the type system can require mandatory configuration steps before build() becomes available.
Conceptually:
Builder<MissingName>
↓
Builder<HasName>
↓
Builder<Complete>
↓
Product
The compiler effectively becomes a checklist.
The Trade-Off
Typestate is powerful, but it is not automatically better.
Too many states can make an API difficult to understand.
For example:
Connection<
Authenticated<
Encrypted<
Connected<
Configured
>
>
>
>
This can become hard to read and maintain.
The goal is not:
“Put everything into the type system.”
The goal is:
Put important invariants into the type system when doing so makes the API clearer and safer.
A good typestate API should reduce cognitive load, not increase it.
A Practical Rule
A useful rule for Rust API design is:
If a state is important enough to check repeatedly,
consider making it part of the type.
For example, compare:
if !connection.is_connected() {
return Err(...);
}
connection.send(data)?;
with an API where send() simply cannot exist on the disconnected type.
The first approach asks every caller to remember the invariant.
The second approach centralizes the invariant in the API design.
The Deeper Rust Lesson
Typestate reveals something interesting about Rust.
Rust is not only about memory safety.
It can also be used to express protocol safety.
Ownership can express who controls a value.
Lifetimes can express how long references remain valid.
Traits can express capabilities and behavior.
Enums can express explicit alternatives.
And typestate can express valid stages in an object's lifecycle.
These mechanisms can be combined.
That is where Rust becomes particularly interesting as a systems language:
Memory model
+
Ownership
+
Lifetimes
+
Traits
+
Type-level state
=
Stronger APIs
The compiler does not merely check whether your syntax is valid.
With thoughtful API design, it can check whether the way you are using a component makes sense.
Final Thought
One of the biggest shifts in Rust programming is learning to stop asking:
“How do I check that this state is valid?”
and start asking:
“Can I design the API so this invalid state cannot be used in the first place?”
That is the heart of typestate.
A boolean can tell you that something is connected.
A type can make it impossible to call send() unless the connection is represented as connected.
That difference may look small in a code example.
In a large codebase, it can become a powerful design boundary.
Rust becomes especially interesting when the compiler is not just checking your implementation—it is enforcing the architecture of your API.