ParallelIQ
Pricing

piqc is free. The platform scales with your fleet.

No published number would fit a 50-GPU team and a 5,000-GPU inference provider equally well. Here's what's in each tier — talk to us for a number that fits yours.

Open source

piqc — $0, forever

Run it against any Kubernetes cluster with a single command and get a full GPU waste report — no account, no payment, no call required. This is the same scanner the platform runs on, not a crippled trial.

Start for Free

The platform

Introspect

Your cluster's X-ray.

Understand the relationship between every model in your fleet and the hardware it runs on. Starts as a read-only one-time scan — graduates to a lightweight agent for continuous monitoring with no changes to your serving stack.

  • Auto-discovers vLLM, Triton, KServe, SGLang, Ollama, TGI
  • Understands model-hardware fit — not just GPU utilization
  • Continuous safety signals: KV cache pressure, OOM risk, queue depth — every 15 seconds
  • Performance and structural signals collected at longer intervals — zero added load on workloads

Remediate

Detect. Recommend. Fix — with a human in the loop.

Every recommendation shows its dollar impact before you approve it. Every action requires operator sign-off. Every change is reversible. The opposite of a black-box autoscaler.

  • Dollar-impact framing per finding — margin recovery for providers, cost reduction for teams
  • Blast radius preview before any change executes
  • Operator approval queue with role-based controls
  • Immutable audit trail — every decision logged under a named operator

Fleet

One optimization layer across every cluster you operate.

Your infrastructure and your customers' — in one view. Whether you run a hosted model API across your own nodes or a deployment platform spanning dozens of customer clusters, Fleet gives you unified visibility and control without replacing anything you already operate.

  • Multi-cluster, multi-cloud, multi-region in one view
  • Covers your clusters and your customers' clusters equally
  • Findings attributed to workload, customer, and cluster automatically
  • Cost and capacity intelligence across the full fleet

Every tier, including piqc alone

  • Nothing executes without explicit human approval
  • Full audit trail — every finding, decision, and action logged under a named operator
  • Deployed entirely inside your environment — model weights, inference data, and customer traffic never leave it
  • Evaluated by a deterministic rules engine, not a black-box model

Questions

Why don't you list prices?

Fleet size, GPU mix, cluster topology, and whether you need on-prem/air-gapped deployment all move the number enough that a single published price would be misleading more often than it would be helpful. Tell us about your fleet and we'll give you a real number quickly — not a multi-week sales process.

Is piqc really free, forever?

Yes. piqc is open source and free to run against any Kubernetes cluster — a single command gets you a full GPU waste report, no account or payment required. It stays free regardless of which Paralleliq platform tier, if any, you adopt later.

Do I have to buy the whole platform, or can I start with just Introspect?

You can adopt Introspect on its own for continuous visibility with no remediation workflow, and add Remediate later once you're ready for recommendations to route through an approval queue. Fleet is additive on top of either, for teams managing multiple clusters — including customers' clusters.

What does a design partner get?

Teams running production GPU clusters who want early access get direct engineering collaboration and preferred pricing in exchange for shaping the roadmap. Mention it when you reach out.

Get more from the cluster you already have.

Start for Free