ParallelIQ
Pricing

piqc is free. The platform scales with your fleet.

No published number would fit a 50-GPU team and a 5,000-GPU inference provider equally well. Here's what's in each tier — talk to us for a number that fits yours.

Open source

piqc — $0, forever

Run it against any Kubernetes cluster with a single command and get a full GPU waste report — no account, no payment, no call required. This is the same scanner the platform runs on, not a crippled trial.

Start for Free

The platform

Step 1

Detect

Your cluster's X-ray.

Understand the relationship between every model in your fleet and the hardware it runs on. Starts with piqc, our open-source scanner — a read-only one-time scan, no signup required. Graduates to a lightweight agent for continuous monitoring with no changes to your serving stack.

  • Auto-discovers vLLM, Triton, KServe, SGLang, Ollama, TGI
  • Understands model-hardware fit — not just GPU utilization
  • Continuous safety signals: KV cache pressure, OOM risk, queue depth — every 15 seconds
  • Performance and structural signals collected at longer intervals — zero added load on workloads
Step 2

Decide & Fix

Detect. Recommend. Fix — with a human in the loop.

Every finding Detect surfaces becomes a recommendation here, with its dollar impact shown before you approve it. Every action requires operator sign-off. Every change is reversible. The opposite of a black-box autoscaler.

  • Dollar-impact framing per finding — margin recovery for providers, cost reduction for teams
  • Blast radius preview before any change executes
  • Operator approval queue with role-based controls
  • Immutable audit trail — every decision logged under a named operator
Step 3

Fleet Scale

The same detect-and-fix loop, across every cluster you operate.

Not a separate product — Detect and Decide & Fix, running across your infrastructure and your customers' infrastructure in one view. Whether you run a hosted model API across your own nodes or a deployment platform spanning dozens of customer clusters, nothing here replaces what you already operate.

  • Multi-cluster, multi-cloud, multi-region in one view
  • Covers your clusters and your customers' clusters equally
  • Findings attributed to workload, customer, and cluster automatically
  • Cost and capacity intelligence across the full fleet

Every tier, including piqc alone

  • Nothing executes without explicit human approval
  • Full audit trail — every finding, decision, and action logged under a named operator
  • Deployed entirely inside your environment — model weights, inference data, and customer traffic never leave it
  • Evaluated by a deterministic rules engine, not a black-box model

Questions

Why don't you list prices?

Fleet size, GPU mix, cluster topology, and whether you need on-prem/air-gapped deployment all move the number enough that a single published price would be misleading more often than it would be helpful. Tell us about your fleet and we'll give you a real number quickly — not a multi-week sales process.

Is piqc really free, forever?

Yes. piqc is open source and free to run against any Kubernetes cluster — a single command gets you a full GPU waste report, no account or payment required. It stays free regardless of which Paralleliq platform tier, if any, you adopt later.

Do I have to buy the whole platform, or can I start with just detection?

You can adopt Detect on its own for continuous visibility with no remediation workflow, and add Decide & Fix later once you're ready for recommendations to route through an approval queue. Fleet Scale is additive on top of either, for teams managing multiple clusters — including customers' clusters.

What does a design partner get?

Teams running production GPU clusters who want early access get direct engineering collaboration and preferred pricing in exchange for shaping the roadmap. Mention it when you reach out.

Get more from the cluster you already have.

Start for Free