Mavits IT Solutions

Projects

Our Private AI Platform

Client
Mavits IT Solutions
Sector
AI infrastructure
Year
2026
Services
AI & LLM Platform Engineering

The challenge

Anyone can recommend private AI. Few can show you theirs. Before we offered private LLM platforms to clients, we set ourselves the same brief a customer would give us: run capable language models on hardware we own, with real production discipline — not a lab setup that falls over when nobody is watching. If we couldn’t operate our own platform to that standard, we had no business building one for anyone else.

What we built

Our platform is a 4-node NVIDIA DGX Spark GPU cluster running Kubernetes (RKE2). Everything above the metal is engineered the way we build for clients:

  • GitOps deployments. Every workload is declared in Git and reconciled to the cluster automatically with Argo CD. There is no “someone changed something on a server” — the repository is the record, and drift corrects itself.
  • Observability. Metrics, logs, and dashboards across nodes, GPUs, and workloads, so operational questions are answered with data instead of guesses.
  • Multi-tenant workloads. Isolated tenant environments let experiments, client-facing services, and internal tools share the hardware without stepping on each other.

On that foundation we did the model work itself: we built, quantized, and now serve our own Mixture-of-Experts model lineage — Vinicius-35B-A3B-v1 — with vLLM, inside our own network, with access control on the serving endpoint. Infrastructure, platform, and model: one team, end to end.

The outcome

The cluster runs production workloads every day, and it doubles as the reference architecture for what we deliver: the same Kubernetes foundation, the same GitOps discipline, the same observability stack. When we size a private AI platform for a client, we are describing hardware and software we already operate — including custom model builds and serving. We don’t sell private AI from a slide deck. Our own platform is our proof.

Outcomes

NVIDIA DGX Spark GPU cluster in production
4 nodes
our own MoE model — built, quantized, and served in-house
Vinicius-35B-A3B-v1
every deployment declared in Git and reconciled automatically
GitOps

Ready to talk about a project?

Tell us what you want to build. We’ll respond quickly with a clear next step — no pitch deck.

Book a consultation