Technology · Compute-Efficiency Layer
July 31, 2026 7 min read Palo Alto, Calif.

One AI server does the work of nearly three — using up to half the energy

A compute-efficiency layer called HXS lets a single NVIDIA H100 carry nearly three times the AI traffic at the same speed, while drawing up to about half the power — without changing the model's answers. It is being offered to one partner.

HXS, a compute-efficiency layer from Trust Carbon Infrastructure (operated by Zenith Flow Innovations LLC), lets one NVIDIA H100 handle nearly three times the AI traffic at the same speed on up to about half the energy — without changing the model's answers. It is being offered to a single partner.

+167%
More capacity
~50%
Less energy
15°C
Up to cooler
100%
Identical output

The short version

The figures come from HXS, a compute-efficiency layer developed by the team behind Trust Carbon Infrastructure. Every number was measured on one NVIDIA H100 80GB running the vLLM 0.26.0 serving stack, comparing the layer on versus off at identical load. The premise is narrow and practical: make the chips a company already owns do more work with less power, without touching the outputs.

Tested on five of the world's largest AI labs

The same result held on the flagship open models of five of the world's largest AI labs — Alibaba, Meta, DeepSeek, Google, and Microsoft. Energy per card and peak temperature, with the layer on versus off:

AI labModel testedEnergy / cardPeak temp
Alibaba Qwen2.5-72B-Instructlarge−26.9%60.947.5 °C
Meta Llama 3.3 70Blarge−27.4%62.447.4 °C
DeepSeek R1-Distill 70Blarge−24.8%61.949.3 °C
Google Gemma 3 27Bmid−17.2%52.643.1 °C
Microsoft Phi-4 (14B)small−25.5%51.040.2 °C

The models are openly published by their labs; nothing here implies endorsement, partnership, or funding by any of them.

Same answers. Less energy. More capacity.

Energy per card — layer on vs off

Reduction measured per model · lower is better · up to −27.4%

MetaLlama 3.3 70B
−27.4%
AlibabaQwen2.5-72B
−26.9%
MicrosoftPhi-4 (14B)
−25.5%
DeepSeekR1-Distill 70B
−24.8%
GoogleGemma 3 27B
−17.2%
0%25%50%

Capacity, one card

Requests/sec, same latency · +167%

Off
109.8
On
293.7

Nearly three times the traffic, no latency percentile worse — up to 62% fewer servers.

Board power (Phi-4, deeper)

Watts per card · −50.5% energy

Off
306 W
On
152 W

Board power roughly halved; peak temp fell 48.2 → 39.2 °C.

Peak temperature — off → on

Cooler chips lower cooling cost and extend hardware life · up to 15 °C

MetaLlama 3.3 70B
−15.0
AlibabaQwen2.5-72B
−13.4
DeepSeekR1-Distill 70B
−12.6
MicrosoftPhi-4 (14B)
−10.8
GoogleGemma 3 27B
−9.5
layer off (hotter) layer on (cooler)scale 38–64 °C

Over about 19 minutes on one server, starting from a draw of 482.7 watts, the energy reduction rose from 38.0% to 50.5% and then leveled off. The energy figure moves with the workload, and the 50.5% figure and the largest capacity gain were measured on one model, Microsoft's Phi-4.

Why it matters: a capacity plan usually buys one chip per unit of demand, and a data center pays for every watt twice — once to power the chip, once to cool it. Here both bills fall together, and the order of chips shrinks, without a single answer changing.

Brayon Michael Pieske, founder
“A capacity plan usually buys one chip per unit of demand. Here the order shrinks, and the power bill with it — and the busiest hour stops being the most expensive.”
Brayon Michael Pieske — founder, Trust Carbon Infrastructure

Bigger than AI — and offered to one company

HXS is not limited to AI serving. The same efficiency layer applies to servers, edge devices, smartphones, IoT hardware, and other computing workloads more broadly; AI serving is simply where it was measured and signed first. That broader reach is a capability — only the AI-serving results here are signed and measured, and no figure is attached to servers, phones, IoT, or other devices. HXS needs no new hardware to adopt.

It is being offered to a single partner — an exclusive license, potentially an acquisition. The technology will be presented in Silicon Valley in the first week of August. A foundational efficiency given to two rivals helps neither, so only one company gets it.

How it can be verified

The AI-serving measurements are cryptographically signed and independently timestamped; qualified parties can verify the complete files under agreement. Results depend on each operation's traffic; validate on your own workload.

The signature is self-signed; only the timestamp is independent (RFC 3161 and OpenTimestamps / Bitcoin). Only the AI-serving results are signed and measured.

See the full proof at hexstellar.com

HXS is a separate technology, built by the team behind Trust Carbon. The complete benchmark, the signed files, and access requests live on the Hexstellar site.

Go to hexstellar.com

About HXS. HXS is a compute-efficiency layer developed by the team behind Trust Carbon Infrastructure, the carbon-verification platform named one of five global winners of the 2025 DPI for People and Planet Innovation Challenge, out of 540 startups from 73 countries. HXS is a separate technology, measured on AI serving; it does not verify carbon. The models named are openly published by their labs; no endorsement or partnership is implied.

More & verification: hexstellar.com · Test platform: NVIDIA H100 80GB · vLLM 0.26.0 · Market reference: H100 cards rent at roughly $1.49–$3.43 per card per hour · Media contact: [email protected]

HXS is developed by the team behind Trust Carbon Infrastructure, operated by Zenith Flow Innovations LLC. HXS is patent-pending. Based in California and Santa Catarina, Brazil.