Distributed Systems Consultancy
For when the cluster is doing something nobody expected
Most systems today are distributed whether they were designed to be or not: a Kafka cluster here, a Cassandra ring there, a dozen services on Kubernetes. Each piece has failure modes the documentation mentions once, in a footnote.
We have run these platforms in production for years, including under our own USP controller. If yours is misbehaving, or you are about to design one, we can help.
Contact usPlatforms we have run in anger
Databases
Apache Cassandra, Apache Ignite, Amazon DynamoDB
Message brokers
Apache Kafka, Apache ActiveMQ Artemis, Apache ActiveMQ, RabbitMQ
Infrastructure and coordination
Kubernetes, Apache Mesos, ZooKeeper, etcd, Consul
In-memory data and compute grids
Hazelcast, VMware GemFire, Redis
Stateful microservices
Akka actors, Axon Framework
What we do with them
Design it with you
Which data store, which broker, where consistency matters and where it does not. Chosen for your workload, not because a conference talk made it sound good.
Sit with the team while it is built
Reviews, pairing, and being there when the first cluster goes up. The design survives contact with reality better that way.
Find the bottleneck
Hot partitions, consumer lag, GC pauses, a coordination service doing more than it should. We measure, we find it, we write up the fix.
Ask what happens when it breaks
Node loss, network partition, a full disk on one replica. We walk through each one with you and fix the ones where the answer is "we are not sure".
Typical engagements
A one-week architecture review with a written report. A two-day incident investigation. A few days a month of design support for a team building its first Kafka or Cassandra system. We are flexible; the constant is that you get something in writing.
Write to us and describe the system. If what you need is a team to build it, see distributed systems development.