Rebuilding our Git storage for agent-scale development

Agents push more often, clone more often and fetch more often than people. We are rebuilding the storage layer while it keeps serving every request.
Repository activity climbs sharply at the far end of the distribution. The busiest repository on the platform saw roughly a billion requests last month.
Where today's architecture meets new demands
Each repository is stored as three replicas. Writes go to all three; reads go to whichever is least busy. That design assumed humans: a few pushes an hour, many reads.
The approach
Minimise coordination
A write should coordinate with as few machines as it possibly can.
func (r *Replica) Apply(ctx context.Context, u Update) error {
if err := r.journal.Append(ctx, u); err != nil {
return fmt.Errorf("journal: %w", err)
}
return r.notifyPeers(ctx, u.Ref)
}Decouple storage from compute
Layer | Before | After |
|---|---|---|
Object storage | local disk | shared, tiered |
Ref updates | 3-phase | journaled |
Reads | any replica | nearest cache |
What comes next
We will publish the benchmarks and the failure-injection results in the next post in this series.