In Development · Pilot-Gated

InferaStack Green Compute Network

A distributed network of energy-aware GPU nodes at grid-connected sites with solar and storage, designed to run as one platform.


“Green” is a rule, not a slogan

A node accepts work only inside the power envelope its site has released, and the envelope state is recorded with every request alongside the region of execution. Base capacity is grid-fed; solar and storage supplement it and provide backup. We do not claim grid-independent operation or a fully renewable supply.

8
GPUs in the R-8 reference residential node
7.5–12 kW
Whole-node draw of the R-8, incl. cooling
240k
Accelerators at 30,000 sites — scenario, not a forecast
0
Nodes validated on hardware yet

What It Is — and What It Is Not

Not another data centre: a network of controllable loads at sites that are already grid-connected and metered, with storage a VPP can already orchestrate.

Five standard nodes

One integrated, sealed and cooled cabinet per site. Two node classes for residential sites with an existing home battery; three for commercial and industrial sites with an existing storage system.

Independent sites, one platform

GPUs and memory are not pooled across sites — there is no cross-site model parallelism. One control plane will manage orders, capacity and telemetry across every site.

Not a data centre

A node is one or two standard OEM servers (4U for the RTX PRO 6000 nodes) in a cabinet. We claim no Tier rating, facility redundancy or whole-machine failover for it. It is designed without a diesel generator, and its closed-loop cooling uses no evaporative water; when the envelope contracts, it curtails and recovers. Its differentiator is that it behaves as a controllable load inside a published power envelope.

Scenario, not forecast

30,000 residential sites × 8 GPUs = 240,000 accelerators at 8–9 kW per site, roughly 240–270 MW. Every one of those figures is a planning scenario until the gates below are passed.


Where a Workload Runs Is Chosen by Data Sensitivity

Three places, one service layer. You choose by data class; the platform is designed to place work accordingly.

Your own environment

Identified personal, clinical or commercially sensitive data. The InferaStack Gateway is designed to deploy privately — your VPC or your own premises — when data must stay in your environment.

A certified colocation facility

Sensitive data that needs certified facility controls. Private GPU environments designed for NEXTDC's Tier IV certified facilities; the certifications are NEXTDC's.

The Green Compute Network

Encrypted, non-identified work only — batch inference, training on de-identified or synthetic data, models over public data — plus overflow when the data class allows.

Placement policy. Residential nodes are designed to accept only encrypted, non-identified work; identified personal or clinical data is routed to your own environment or a certified facility, and the region of execution is recorded on every request, so a council or health service can see where each job ran.

What a Node Will Offer

Reserved capacity for AI workloads that run all day — designed, not yet available.

Reserved Capacity · In Development

Fixed-budget model service

Customers will be able to reserve capacity on a pre-deployed open model so their AI agents can run around the clock within a fixed budget, draining to a hosted backend when a node is curtailed. Dedicated endpoints on one or two GPUs, or shared endpoints with a reserved share plus idle-capacity borrowing, scheduled per request.

  • OpenAI-compatible API through the InferaStack Gateway
  • First model configuration: a 30–35B-parameter open model at 8-bit (FP8) precision on one 96 GB GPU
  • Envelope state and region of execution recorded per request
GPU Capacity · In Development

GPU VMs and managed workers

Isolated Linux VMs on 24 GB, 48 GB or a full 96 GB GPU; two-GPU managed workers within one compute module for larger models; spot and batch on whatever the envelope leaves idle.

  • GPU partitions (MIG) for the 24 GB and 48 GB sizes, a whole GPU (passthrough) for 96 GB
  • Spot and batch are the curtailable classes when the envelope contracts
  • No cross-site pooling — capacity is sold per site

How a Node Behaves

Control priority is fixed by design. The compute layer sits at the bottom of it.

1 — Never overridden

Device protection

Battery management, power conversion, electrical protection and server thermal protection.

2 — Site policy

The owner's rules

Household or business backup reserve and operating limits set by the site owner.

3 — Energy envelope

Envelope publisher (ConnectVPP)

Designed to publish the site's power envelope from live telemetry — load, storage state, solar, network limits, market dispatch windows — as a ceiling in kilowatts plus the state that produced it: grid-only, solar-matched or storage-backed. ConnectVPP is the intended publisher; the interface is being confirmed with it.

4 — Service layer

InferaStack control plane

Admits and schedules work inside the envelope. Firm reserved capacity is sized to the envelope's grid-fed base plus storage-backed reserve; everything else yields to it. Spot, batch and borrowed idle capacity are curtailed first; then GPUs are power-capped through driver-level limits; then, as a last resort under agreed conditions, work drains to a hosted backend. Safe local operation is designed to continue if the cloud is unreachable.


The Energy Side, with ConnectVPP

By design, ConnectVPP publishes the envelope from live site telemetry and the InferaStack control plane obeys it. Storage stays the owner's asset throughout.

1 · Live telemetry in

ConnectVPP's platform is designed to read the live site telemetry and release the permitted envelope to the node — level 3 of the control hierarchy above. ConnectVPP describes that platform as Australia's B2B virtual power plant platform with sub-100 ms orchestration.

2 · Storage stays the owner's

When compute load is low, the owner's storage remains available for dispatch by its VPP operator, including into AEMO's frequency-control markets. A VPP event is designed never to interrupt a contracted continuous service automatically.

Measured, not claimed. Every request records the region of execution and the envelope state it ran under, so solar-matched and storage-backed hours can be reported per site — that per-request record is the operational meaning of “green”. Per site, we intend to report energy per request and the share drawn in solar-matched hours, including solar that would otherwise have been curtailed; cooling water per GPU-hour (none by design — closed-loop, no evaporative use); backup fuel per GPU-hour (none by design — the node curtails and recovers); and envelope compliance. Comparisons with centralised facilities or public cloud will be published as measured results against a stated baseline, not as estimates. Until measured, every such figure is a target.

Status. How ConnectVPP publishes the envelope is to be confirmed; running our envelope logic against a simulated ConnectVPP envelope is a Gate 0 deliverable. The energy-side components are described in the power and storage reference design; what has and has not been tested is tabulated under Where Things Stand below.

The Nodes

Preliminary engineering targets, September 2026. Nothing below has been validated on hardware.

NodeGPUsWhole-node input (range, incl. high-temperature bound)Energy per daySite class
R-88× NVIDIA RTX PRO 6000 Blackwell Server Edition7.5–12 kW192–288 kWhResidential
R-4H4× NVIDIA H200 NVL (one 4-way NVLink domain)4.5–7 kW120–168 kWhResidential
C-1616× RTX PRO 6000 as two 8-GPU compute modules15–24 kW384–576 kWhCommercial & industrial
C-8H8× H200 NVL (two 4-way NVLink domains)7.5–12 kW192–288 kWhCommercial & industrial
C-8S8× H200 SXM on one HGX baseboard10–18 kW288–432 kWhCommercial & industrial
Whole-node means whole node. Energy per day runs from the reference draw (8 kW for an 8-GPU node) to the high-temperature bound over 24 hours at full load. The figures include cooling, network and UPS auxiliaries, not just GPU board power (which is 4.8 kW for the 8-GPU node). Each node needs a dedicated three-phase 400/230 V feed, closed-loop cooling with compressor-based heat rejection sized for a 45 °C day, symmetric 10 Gbps fibre with an independent backup, and a sealed outdoor cabinet. The residential acoustic target of 40 dBA at 1 m is unmeasured. The compute platform is a standard OEM server — 4U for the RTX PRO 6000 nodes; the cooling configuration for the residential product is still to be confirmed with the OEM.

Where a Node Can Live

Residential where the electrical envelope permits; commercial and industrial sites are the likely primary class.

Residential

Three-phase homes with a real battery

A residential node needs a dedicated three-phase feed with headroom, an existing battery whose inverter can back up all three phases and carry the node, deliverable commercial fibre, an independent outdoor equipment area, and a neighbour-noise assessment. A typical Australian house is single-phase and uses roughly 15–20 kWh a day; an 8-GPU node uses 192–288 kWh. What share of the housing stock qualifies is the first thing the site process has to establish.

Commercial & industrial

Sites with existing solar and storage

Premises with a three-phase solar-and-storage system whose power conversion can carry the node as a dedicated backup load. The storage system stays a separate, safety-isolated facility outside the IT cabinet; the node is designed to take only surplus inside your envelope and never to draw on your backup reserve. The 16-GPU and H200 nodes are designed for this class.

Site process

Seven stages, no shortcuts

Authorised referral → remote pre-screen → joint survey (electrical, structural, thermal, noise, network, planning, insurance) → commercial approval → installation and acceptance with measured power → commissioning and monthly operating review → periodic review. Residential installation only after a full node passes acceptance.


The Gated Pathway

Progress by acceptance results and paid reserved demand, not by calendar.

Gate 0

Bench

Four GPUs in a controlled environment: single-GPU model endpoint, MIG-backed GPU partitions, dual-GPU peer-to-peer, and the curtail–drain–recover logic run against a simulated power envelope.

Gate 1

One node

One standard 8-GPU node under a 72-hour continuous load, with power, temperature and throttling recorded. Acceptance targets: p95 time-to-first-token ≤2 s, p95 inter-token ≤50 ms, ≥99.5% request success.

Gate 2

First sites

3–10 accepted sites — only after three tests pass: the whole cabinet at design temperature, noise at 1 m and at the neighbours' boundary, and three-phase battery switchover.

Gates 3–4

Tens, then hundreds

Around 50 sites, then 100–200. Advancement on acceptance results and paid reserved demand, not on the calendar.


Where Things Stand

Tested versus proposed, as of September 2026.

StatusWhat
Built and deployed to our own AWS account, not releasedThe OpenAI-compatible InferaStack Gateway with per-request metering, per-key budgets and metadata-only audit records.
Built and unit-tested, never run on hardwareSelf-hosted-node invocation, energy-envelope admission control, and drain to a hosted backend when a node is curtailed or unreachable.
Design onlyAll five nodes and the node software platform. No site acceptance anywhere; no acoustic, thermal or electrical validation performed.

Our curtail–drain–recover protocol is working code with tests, not a slide. What we have never done is run it against a real node at a real site — which is exactly what Gate 0 and Gate 1 exist to establish.


Register Interest to Host a Node or Reserve Capacity

This program is for site hosts with three-phase supply and storage, for energy partners, and for research groups, SMEs and AI product teams that need fixed-budget or schedulable capacity. Register interest and we will tell you where the program stands.

Register interest →