pub struct Sim<P: Protocol> { /* private fields */ }Expand description
A deterministic run of P across several processes.
Ordering is total and reproducible: the queue is keyed by (time, sequence), so events at
the same virtual instant are processed in the order they were scheduled, on every run.
Implementations§
Source§impl<P> Sim<P>
impl<P> Sim<P>
Sourcepub fn run_scenario(
scenario: &Scenario<P::Cmd>,
build: impl FnOnce(Config, &[NodeId]) -> Sim<P>,
) -> Sim<P>
pub fn run_scenario( scenario: &Scenario<P::Cmd>, build: impl FnOnce(Config, &[NodeId]) -> Sim<P>, ) -> Sim<P>
Execute a description, and return the run it produced.
build is handed the scenario’s configuration and membership and returns a simulator over
them — normally Sim::new(config, nodes, make), plus whichever opt-ins the protocol needs
(deliver_session_events, enable_codec_check). It takes both rather than closing over
them because the shrinker will hand it smaller ones, and a builder that ignored its
arguments would quietly keep running the original.
The clock is advanced to each step’s time and the step applied, exactly as a hand-written
test would: run_until(at) then the call. Two steps at the same time are applied with
nothing dispatched between them, which is what makes “command, then break the session
carrying it” expressible.
Source§impl<P> Sim<P>
impl<P> Sim<P>
Sourcepub fn new(
config: Config,
nodes: &[NodeId],
make: impl FnMut(NodeId) -> P + 'static,
) -> Self
pub fn new( config: Config, nodes: &[NodeId], make: impl FnMut(NodeId) -> P + 'static, ) -> Self
Build a run over nodes, constructing each process with make.
Sourcepub fn delivery_bound(&self) -> Option<Duration>
pub fn delivery_bound(&self) -> Option<Duration>
The upper bound on delivery, when the run is synchronous.
A protocol whose correctness rests on this bound should be configured from it rather than from a timeout that happens to work.
Sourcepub fn trace(&self) -> &ProtoTrace<P>
pub fn trace(&self) -> &ProtoTrace<P>
The record of what has happened so far.
Sourcepub fn record_notes(&mut self)
pub fn record_notes(&mut self)
Record what protocols narrate, so the trace holds their decisions beside what happened.
Off by default: a run pays nothing for an audience it does not have, and the protocol’s own
code is the same either way — it calls Cx::note regardless, and the sink discards. That is
what makes narrating unable to change the run.
Sourcepub fn protocol(&self, node: NodeId) -> Option<&P>
pub fn protocol(&self, node: NodeId) -> Option<&P>
Borrow a process, for inspecting state a trace cannot show.
Sourcepub fn at(&self, node: NodeId) -> &P
pub fn at(&self, node: NodeId) -> &P
Sim::protocol for a process the test knows is running. Panics naming the process
otherwise, which is more use in a failure than unwrap’s line number.
Sourcepub fn storage(&self, node: NodeId) -> Option<&MemStore<P::Meta, P::Entry>>
pub fn storage(&self, node: NodeId) -> Option<&MemStore<P::Meta, P::Entry>>
Borrow what a process has written down. None if it has written nothing.
Sourcepub fn nodes(&self) -> impl Iterator<Item = NodeId> + '_
pub fn nodes(&self) -> impl Iterator<Item = NodeId> + '_
The processes in this run, in a stable order.
Sourcepub fn command(&mut self, node: NodeId, cmd: P::Cmd) -> OpId
pub fn command(&mut self, node: NodeId, cmd: P::Cmd) -> OpId
Hand cmd to node at the current time, and take an identity naming that operation.
The identity is a return value rather than a parameter, so a caller with no interest in it
carries on as before. It names the operation in the trace: Trace::invoked_at says when
the process handled it, and Trace::why_not_begun says why it never did.
Sourcepub fn command_at(&mut self, node: NodeId, after: Duration, cmd: P::Cmd) -> OpId
pub fn command_at(&mut self, node: NodeId, after: Duration, cmd: P::Cmd) -> OpId
Hand cmd to node after after has elapsed, and take an identity naming that operation.
Sourcepub fn crash_on_next_write(&mut self, node: NodeId)
pub fn crash_on_next_write(&mut self, node: NodeId)
Arm the next write by node to be the one it dies inside.
Whether that write landed is decided by the seed, and the process cannot tell: what it reads on recovering is the only evidence.
Sourcepub fn crash(&mut self, node: NodeId)
pub fn crash(&mut self, node: NodeId)
Crash node: it stops handling events and loses everything volatile.
Its protocol state is replaced with a freshly initialised one, and its pending timers and
anything held for it are discarded, so a restart resumes having forgotten what it
delivered. This is what a real process gets. For a stall that preserves state, use
Sim::suspend.
Sourcepub fn suspend(&mut self, node: NodeId)
pub fn suspend(&mut self, node: NodeId)
Suspend node: it stops handling events but keeps its state, its timers, and everything
addressed to it while it is away.
Not what a crash does. This is a stall — the process is descheduled and comes back having missed nothing, because the timers, deliveries and scope events that came due meanwhile were held rather than dropped. Losing them would be losing a message inside a session that never ended, which is the one thing this model forbids of a layer.
For unreachability use Sim::partition, and for failure Sim::crash. Resume with
Sim::resume; Sim::restart is for a crashed process and re-runs its startup branch.
Sourcepub fn resume(&mut self, node: NodeId)
pub fn resume(&mut self, node: NodeId)
Resume a suspended node: everything held while it was away is dispatched now.
No startup branch runs. Nothing was lost, so there is nothing to recover and nothing to
initialise — replaying on_init or on_recovery over intact volatile state would be
telling a process it restarted when it did not. That is what separates this from
Sim::restart.
Note what a resumed process is not told: that time passed. Its clock ran while it could not read it, so anything that measures silence — a failure detector — comes back with stale evidence and a timer due immediately. That is what a stall does to a real process, and it is why the synchronous model excludes one.
Sourcepub fn restart(&mut self, node: NodeId)
pub fn restart(&mut self, node: NodeId)
Restart a crashed node, which takes its startup branch.
A crash discarded its volatile state, its timers, and anything that was on its way, so
there is nothing held to release. What survived is in storage, and is handed back through
on_recovery. For a suspension use Sim::resume.
Sourcepub fn is_stopped(&self, node: NodeId) -> bool
pub fn is_stopped(&self, node: NodeId) -> bool
Whether node is currently stopped, for either reason.
Sourcepub fn partition(&mut self, groups: &[&[NodeId]])
pub fn partition(&mut self, groups: &[&[NodeId]])
Split the network into groups. Messages between groups are not delivered.
The special case of Sim::sever in which the severed pairs are exactly those spanning two
groups — so reachability is transitive here, and a process in a group reaches every other
member. That is the easy case, and it is not the only one; see sever.
Replaces whatever was severed before, so calling this after sever discards that severing.
Sourcepub fn sever(&mut self, a: NodeId, b: NodeId)
pub fn sever(&mut self, a: NodeId, b: NodeId)
Cut a and b off from each other, in both directions, leaving every other pair alone.
This is what a grouping cannot express. Severing one pair of three processes leaves a
bridge: A reaches B and B reaches C, but A does not reach C. All three are
correct, none of them is wrong about what it can see, and there is no group any of them
belongs to — which is the case every layer above that depends on processes agreeing about
who is reachable has never been asked about.
Severing is symmetric. A link that works one way and not the other is a different fault and
a harder question for the session model, which treats a session as a property of a pair; see
this change’s design.md.
Sourcepub fn reconnect(&mut self, a: NodeId, b: NodeId)
pub fn reconnect(&mut self, a: NodeId, b: NodeId)
Restore connectivity between a and b, leaving every other severing in place.
Sourcepub fn reachable(&self, a: NodeId, b: NodeId) -> bool
pub fn reachable(&self, a: NodeId, b: NodeId) -> bool
Whether a and b can currently reach each other.
For asserting the topology a test built rather than assuming it: severing one pair of four processes is not a bridge, and a test that thinks it is would be testing nothing.
Sourcepub fn heal(&mut self)
pub fn heal(&mut self)
Restore full connectivity, discarding every severing however it was made.
Sourcepub fn step_now(&mut self)
pub fn step_now(&mut self)
Process every event scheduled within d of now.
Dispatch everything scheduled for the current instant, and nothing later. The clock does not
move.
For sequencing a test by events rather than by durations: a command is scheduled, not run,
so command(...) followed by break_session(...) breaks a session with nothing in flight.
step_now() between them runs the handler, whose sends then sit in the queue at their
latency — in flight, and the break finds them. The older idiom, run_for(1 ms) with a
comment, depends on the latency being longer than the millisecond.
Sourcepub fn step(&mut self) -> bool
pub fn step(&mut self) -> bool
Dispatch the next scheduled event, moving the clock to it. false when there is nothing
left to dispatch, or the step budget is spent.
For a test searching for a state that one event creates and the next may destroy — “exactly
one process has decided” — stepping by event cannot skip it, where run_for(1 ms) can when
two events fall inside the millisecond.
pub fn run_for(&mut self, d: Duration)
Source§impl<P> Sim<P>
impl<P> Sim<P>
Sourcepub fn enable_codec_check(&mut self)
pub fn enable_codec_check(&mut self)
Round-trip every delivered message through the wire codec.
Off by default: the simulator moves typed values, so a codec defect cannot be mistaken for a protocol defect. Turn it on to check that messages actually survive encoding, without paying for it on every run.
Source§impl<P> Sim<P>
impl<P> Sim<P>
Sourcepub fn enable_tracing(&mut self)
pub fn enable_tracing(&mut self)
Emit every recorded event to whatever tracing subscriber is installed, as it is recorded.
Off by default, like the codec check: a run pays nothing for an audience it does not have. Turning it on does not change the run — the events are the same ones the trace already holds, in the same order, and nothing a protocol can observe is affected.
Pair it with Sim::record_notes to see what the protocols said as well as what happened
to them; without it the rendering shows the run, which is what the trace showed before this
existed.
Source§impl<P> Sim<P>
impl<P> Sim<P>
Sourcepub fn session_epoch(&self, a: NodeId, b: NodeId) -> Option<u64>
pub fn session_epoch(&self, a: NodeId, b: NodeId) -> Option<u64>
The current epoch of the session between a and b, if one is established.
Sourcepub fn break_session(&mut self, a: NodeId, b: NodeId)
pub fn break_session(&mut self, a: NodeId, b: NodeId)
End the session between a and b, discarding an unknown suffix of what was in flight.
A new session opens at a higher epoch the next time either sends and the pair is able to communicate.
Sourcepub fn has_session(&self, a: NodeId, b: NodeId) -> bool
pub fn has_session(&self, a: NodeId, b: NodeId) -> bool
Whether a session currently exists between a and b.
Source§impl<P> Sim<P>
impl<P> Sim<P>
Sourcepub fn deliver_session_events(&mut self)
pub fn deliver_session_events(&mut self)
Deliver session events to the protocol as scope events.
Opt-in, like the codec check: a protocol that declares no scopes cannot receive one, and the bound lives only on this method so ordinary runs need nothing.
Forgetting this is silent, and it disables everything the session layers do. Sessions
still open, end and lose their suffixes; no layer is ever told, so every resend clause is
dead, every [session] tag is unearned, and nothing fails to say so. A session-based run
of a protocol whose Scope is inhabited should call it, and
forgetting_deliver_session_events_silently_disables_the_whole_bridge is what that costs.