The Fleet is a household lab of models and agents: a dozen local models served across Mac and DGX-class machines, behind one endpoint.
A smart router decides where each request runs. Memory-aware scheduling accounts for what parameter counts don't tell you, because how much memory a model actually needs depends on its architecture, not just its size.
Health dashboards watch the machines, and agents maintain the fleet they run on. It is infrastructure that operates itself, observed rather than babysat.
The Fleet exists because the questions that matter about running AI can't be answered from slideware. It is the standing evidence behind everything this site claims.