ParallelIQ
AI Infrastructure

What is Paralleliq?

Common questions about Paralleliq, piqc, and how the two fit together.

What does Paralleliq do?

Paralleliq makes your GPU fleet cheaper to run and safer to operate. It detects both waste — idle capacity, tier misplacement, CPU:GPU imbalance — and risk, like OOM risk, across Kubernetes-based inference clusters. Waste findings come with a quantified dollar impact; risk findings are flagged with severity and evidence rather than an invented cost. Either way, findings surface as operator-approved recommendations with a full audit trail.

The result is a lower $/token cost on capacity you already have, fewer outages from misconfigured deployments, and a clear signal on when you actually need to order more GPUs — rather than ordering early because you're not sure.

Paralleliq is not affiliated with Parallel, a separate company building AI agent infrastructure — the names are unrelated.


What is piqc?

piqc is Paralleliq's GPU waste scanner for Kubernetes AI inference clusters. It's a single command that runs read-only against a cluster — nothing is installed permanently — and reports idle GPUs, tier misplacement, OOM risk, and CPU:GPU imbalance, each with an estimated dollar impact.


Who makes piqc and is it open source?

piqc is built and maintained by Paralleliq. Its source is publicly available on GitHub under the Business Source License 1.1 (BUSL) — you can read it, run it, and modify it for non-production use; the license converts to Apache 2.0 in January 2028. It is source-available rather than OSI-approved open source in the interim.


What is the piqc GPU scanner used for?

Teams run piqc to get a free, point-in-time read on GPU waste and risk before adopting anything further — it scans every workload in a cluster and reports which models are on the wrong GPU tier, which nodes have idle or unallocated capacity, and where CPU:GPU ratios are starving GPUs of work, each with a per-workload cost estimate — plus which pods are at OOM risk, flagged by severity rather than a dollar figure.


How does Paralleliq detect GPU waste?

Paralleliq combines Kubernetes resource state with model-level context — what model is running, what hardware tier it actually needs, how it behaves under load — to find waste and risk that pure metrics monitoring misses. Waste findings — tier misplacement, idle or dark capacity, CPU:GPU imbalance — are each detected with a quantified dollar impact, not just a utilization percentage. Risk findings like OOM risk are flagged with severity and evidence instead, since the cost of an outage isn't something we invent a number for.


How does Paralleliq help reduce $/token costs?

Most $/token cost comes from GPU capacity you're paying for but not using efficiently — models on the wrong tier, idle nodes, OOM-driven over-provisioning. Paralleliq finds those gaps, quantifies their dollar impact, and routes operator-approved fixes through a full audit trail.

piqc is the starting point: a free, read-only scan that shows exactly where waste is. The full platform builds on that to deliver ongoing recommendations across a fleet. Read more about how the optimization layer works.


Does Paralleliq use AI to decide what to do with my infrastructure?

No. Findings come from a deterministic rules engine evaluated against observed telemetry, not a model — the same input always produces the same output, and every finding traces to a readable rule. On top of that, no recommendation executes automatically: every action requires explicit human approval first. The combination — explainable rules plus a human in the loop — means there is no autonomous AI agent making changes to your infrastructure.


Does Paralleliq work with on-premises GPU clusters?

Yes. Paralleliq works wherever Kubernetes runs — on-prem data centers, private cloud, co-location facilities, and sovereign AI environments are all supported. There is no dependency on a public cloud provider. piqc connects to your cluster via standard kubeconfig access, and the full platform is deployable entirely within your own infrastructure if data residency or air-gap requirements demand it.


Does piqc install anything permanently in our cluster?

No. piqc is a single read-only command — it queries your cluster's existing Kubernetes API and exits. Nothing is deployed, no agent is left running, no cluster state is modified. It requires only read access to pod and node resources, which you can grant via a temporary kubeconfig or a scoped service account that you revoke after the scan.


What does it cost to get started with Paralleliq?

piqc is free. You can run it against any Kubernetes cluster today with a single command and get a full GPU waste report at no cost. The full Paralleliq platform — ongoing recommendations, fleet-wide monitoring, approval workflows, and audit trail — is available under a commercial subscription. Contact us to discuss pricing based on your fleet size.

Try piqc on your own cluster

One read-only command, no installation, a free report of GPU waste across your fleet.

Get more from the cluster you already have.

Start for Free