Artificial Intelligence
NVIDIA and Palantir Deploy Sovereign AI Stack for Supply Chains

Palantir Technologies (PLTR ) and NVIDIA on September 10, 2026, announced a collaboration to bring sovereign AI to critical supply chains, with the AI stack first being deployed within NVIDIA’s own supply chain operations. The deployment creates an AI stack that brings NVIDIA Nemotron open models into Palantir Foundry and the Artificial Intelligence Platform (AIP), grounded in the Palantir Ontology.
According to the announcement, the stack aims to create supply chain visibility, identify constraints, continuously codify operational expertise and guide decisions at machine speed while the organization retains control and ownership of its proprietary data. Within Palantir AIP, NVIDIA cuOpt software supports optimization and scenario planning, enabling teams to model supply constraints, assess tradeoffs and understand the operational impact of allocation decisions. Post-trained Nemotron open models recommend actions, explain tradeoffs and flag emerging risks, while supply chain experts retain control of final decisions.
NVIDIA states that it operates one of the most complex supply chains, spanning millions of parts, thousands of suppliers and a global network of manufacturing partners. Bringing a rack-scale AI system to production requires the coordinated availability of compute, memory, networking, power, cooling and mechanical components, and the company states it works closely with its ecosystem to secure supply for the 1.3 million parts in each Vera Rubin rack.
The NVIDIA Deployment From Wafer-Out to First Token
An NVIDIA Technical Blog post, also dated September 10, 2026, and written by NVIDIA solutions architects Nell Barber, Rana Haber and Aastha Jhunjhunwala, details the deployment’s mechanics. NVIDIA measures supply chain performance from wafer-out to first token: time-to-rack runs from silicon leaving the fab to an assembled system arriving on a data center floor, and time-to-token covers power, cooling, networking and the software stack that makes the infrastructure productive. One Grace Blackwell NVL72 compute tray, one of eighteen in a single rack, requires two Grace CPUs, four Blackwell GPUs and thirty-two HBM3e stacks. The supply chain created for Vera Rubin is twice as large as Grace Blackwell’s, according to the post.
Contract manufacturers cannot start assembly until every component has arrived from one of three pools: parts from NVIDIA directly, parts NVIDIA stocks on consignment, and parts from suppliers. NVIDIA runs a Time of Ownership clock from the moment a manufacturing site receives material until it leaves as part of a sub-assembly or product. Each week it must decide what and how much material to allocate to each manufacturing site, a critical material allocation problem that planners manually rework weekly, with the allocation running through the current quarter and the next.
The NVIDIA supply chain operations team worked with Palantir to build a Digital Supply Chain Intelligence command center on Palantir Foundry, where the Ontology connects materials, manufacturing sites, commits, capacity, allocations, production outputs and unstructured qualitative signals into one governed data layer. NVIDIA cuOpt, an open-source library for GPU-accelerated decision optimization, draws its inputs from the Ontology and solves the weekly allocation as a mixed-integer linear program whose objective minimizes Time of Ownership. The solver also reports which constraints are binding, and because the solve is fast, planners can explore scenarios around the answer, such as ten percent less memory in a period or a new manufacturing site coming online.
Back-testing historical allocation decisions against actual outcomes showed that human planners outperformed the solver by working from information it could not see: emails exchanged with partners that week, severe weather forecasts for key regions, ongoing geopolitical events, supplier debrief transcripts and years of accumulated expertise. The workflow therefore captures each allocation decision, the rationale behind it, the expected result and the actual outcome.
NVIDIA post-trained the open-weight Nemotron 3.5 Lightning model, which has 30 billion parameters with roughly 3 billion active per forward pass, on the captured decisions. The pipeline uses NeMo Anonymizer to remove personally identifiable information and obfuscate sensitive fields, NeMo Data Designer to expand and balance examples with synthetic data, and NeMo AutoModel to train a small set of LoRA adapter parameters while the base weights stay frozen. Palantir Autopilot manages the lifecycle end to end, launching each job from Ontology data, monitoring the deployed model and keeping lineage intact from data to model version to recommendation.
Reported Benchmark Results
On NVIDIA’s development benchmark, the post-trained Lightning model reached 86.7% allocation-decision accuracy, compared with 55.5% for the larger Nemotron 3 Ultra and 17.5% for the base Lightning model, the blog reports. On metrics that weight decision types equally rather than examples, the post-trained model scored 58.6% balanced accuracy against Ultra’s 42.0%, and 57.5% Macro-F1 against Ultra’s 39.5%. The LoRA training run finished on two NVIDIA B200 GPUs in minutes. The blog states that the model’s gains are concentrated in the domain it was post-trained on, and that future production risk forecasting remained difficult despite fine tuning.
Accepted, edited and overridden recommendations are written back into the Ontology, accumulating until there is enough representative data to justify another governed training run. The companies plan to use that feedback for reinforcement learning, with accepted and overridden recommendations forming preference pairs covering allocation correctness, policy compliance and evidence grounding; the model never retrains itself in production.
The deployment runs on NVIDIA reference architectures and the jointly developed Palantir Sovereign AI Operating System Reference Architecture (SAIOS), which is supported by Dell Technologies (DELL ) and Cisco. The AI stack can be deployed on premises with system manufacturers including Cisco and Dell, or in co-location and cloud environments with Rackspace and Nebius.
“NVIDIA has arguably the most valuable, intricate and complex supply chain in the world. Our sovereign stack, powered by Nemotron models and Ontology, is delivering capabilities that exceed the frontier while providing alpha protection qualities unavailable otherwise,” said Alex Karp, cofounder and CEO of Palantir Technologies.
“Supply chains are the operating system of the physical economy, and AI factories are among the most complex systems ever built,” said Jensen Huang, founder and CEO of NVIDIA.
The companies plan to extend learnings from NVIDIA’s deployment to enterprises across manufacturing, energy, healthcare, automotive and aerospace, and the announcement cites agriculture, manufacturing, pharmaceutical, retail, technology and government organizations as able to use the AI stack to optimize their own supply chain operations. The AI stack and its applications across industries will be showcased at Palantir’s AIPCon 11 conference.












