Decoding Jensen Huang’s “five-layer cake” — from power, chips, and data centers to models and applications — and the leadership decisions the AI era demands
Author: Loi Doan (Luke) · SpeedUP Technology Vietnam · 25+ years in IT infrastructure, cloud, and AI · 15+ years at FPT Corporation
When people talk about AI, most of us picture ChatGPT, chatbots, or the “smart” applications users see on screen. But the application is only the top layer. At the World Economic Forum in Davos in January 2026, Jensen Huang — founder and CEO of NVIDIA — described AI as a “five-layer cake”: power, chips, data-center/cloud infrastructure, models, and applications. That framing takes the AI conversation out of the narrow scope of a software project: to get good AI, an organization needs a sufficiently strong, stable, and secure chain of infrastructure underneath it.
In a conversation with BlackRock CEO Larry Fink at Davos 2026, Jensen Huang called the current AI wave “the largest infrastructure buildout in human history.” What stands out is not just the scale of capital involved. More importantly, AI is increasingly viewed as a foundational capability of the economy — comparable to electricity, telecommunications, or digital infrastructure — rather than just another piece of software bolted onto a business.
The difference lies in the nature of the computational load itself. Traditional software systems mainly execute pre-programmed rules. The new generation of AI has to process huge volumes of text, images, voice, and unstructured data while running enormous amounts of parallel computation. So behind an answer that appears in a few seconds sits an entire chain of power, cooling, chips, networking, storage, and operating software that must work together almost continuously. The more AI is used, the more pressure builds on the layers underneath it.
From a national perspective, Jensen Huang also warned that a country that fails to build this foundation — starting with energy and compute capacity — risks becoming nothing more than a buyer of AI services. For Vietnam, this is a long-term story about competitiveness, data sovereignty, and technological self-reliance.
This is also the point many businesses tend to underestimate. Leaders see the chatbot, the virtual assistant, the automated workflow; but when the AI system is slow, expensive, or unstable, the real cause is usually deeper: storage that cannot feed data fast enough, a congested network, GPUs left waiting, a cooling system hitting its limit, or a security gap between the cloud and on-premise infrastructure. So the right question is not only “which AI application should we use?” but also “which infrastructure will keep that application running stably, safely, and cost-effectively?”
AI can be pictured as a five-layer cake, where each upper layer is only as stable as the layer beneath it:

Figure 1. The five layers of AI infrastructure per Jensen Huang’s framework — illustration by SpeedUP.
| Layer | Name | What it is | Where value / risk shows up |
|---|---|---|---|
| 5 | Application | Virtual assistants, process automation, predictive analytics, request routing | Business value — the tangible results leaders see directly |
| 4 | Model | Operational and physical AI, digital twins, cybersecurity automation | Intelligence — quality depends on data and operations |
| 3 | Infrastructure / AI Factory | Storage, internal networking, orchestration, monitoring, security telemetry | Where small latency gets amplified into major incidents |
| 2 | Accelerated Computing | GPU clusters, high-speed memory, ultra-fast interconnect fabric | Compute economics — bounded by memory and bandwidth |
| 1 | Power | Building capacity, cooling, backup systems | The physical foundation — where failures begin |
Table 1. The five layers of the “cake” per Jensen Huang, mapped against operational reality.
The most important principle for operators: the place where a fault is reported and the place that actually causes it are usually two different places. “The training server is running slow” or “the model has gotten worse” sound like software bugs, but the real culprit is often the power layer, the network layer, or the storage layer. A few quick examples:
That is the real value of the “five-layer” model: it forces us to see AI as an end-to-end system. For business leaders, this view helps avoid a costly mistake — pouring heavy investment into the application layer while the foundation layers underneath are not yet ready.
Power is the bottom layer and also the hardest barrier to get around. As engineers like to say, “there’s no shortcut around the power problem” — capacity and cooling capability are already locked in long before the first GPU is racked. According to the International Energy Agency (IEA), global data-center electricity consumption is forecast to nearly double, from roughly 415 billion kWh in 2024 to around 945 billion kWh by 2030 — close to Japan’s entire current electricity consumption.
The AI-specific share is growing far faster: AI-dedicated servers are increasing by about 30% a year, versus roughly 9% for ordinary servers. The US and China together account for nearly 80% of the global increase. In the US, by 2030 data centers are projected to consume more electricity than the aluminum, steel, cement, and chemical industries combined.

Figure 2. Global data-center power consumption is set to double by 2030 (source: IEA).
Power has therefore stopped being a matter for the building-facilities team alone and has become a direct constraint on expansion. Goldman Sachs estimates the US will face a shortfall of about 9.3 GW of power capacity in 2026, potentially rising to 45 GW by 2028 — large enough that major tech companies have already signed contracts for nuclear power or small modular reactors (SMRs) to secure supply.
The biggest difference lies in power density per rack. A typical enterprise or data-center rack draws 4–8 kW; a rack of modern GPU servers commonly exceeds 50 kW; and NVIDIA’s latest GB200 NVL72 design reaches roughly 132 kW per rack — 16 to 30 times higher. That figure largely determines which cooling method is even feasible:

Figure 3. The cooling threshold by rack power density — beyond roughly 40 kW, air is no longer enough.
At the same time, Power Usage Effectiveness (PUE) — the metric for a data center’s power efficiency — is being pushed down, from a typical 1.5–1.8 PUE to 1.1–1.2 PUE with liquid-cooled designs. The closer to 1.0, the less power is wasted on cooling and on the data center overall. In other words, once the load exceeds roughly 40 kW per rack — a common threshold for most high-density GPU clusters — liquid or immersion cooling becomes close to a technical necessity, especially in hot, humid climates.
There is a classic example of the “where the fault is reported differs from where it is caused” principle: training time for an AI system slowed by about 12% every afternoon and then recovered on its own overnight. The engineering team suspected a bug in the data-processing software. In reality, when the weather got hot and the building’s air-cooling system became overloaded in the afternoon, the GPU servers automatically throttled down to avoid overheating — the computation itself was identical, it just did less work per second. The “fix” was in the cooling system, not the software. The lesson: whenever the performance of an entire AI infrastructure stack rises and falls with the time of day, it is almost always a hardware and cooling issue — software doesn’t know what time it is.
If air cooling has been the default data-center architecture for decades, AI is now forcing the industry to revisit that assumption. As power density rises from a few kW to tens, even over 100 kW per rack, the question is no longer simply “how much colder should we make the room,” but how to pull heat away from the chip fast enough for the GPU to sustain its designed performance. The Uptime Institute notes that AI is accelerating adoption of both cold-plate/direct-liquid cooling and immersion cooling, though actual deployment still depends on how well it integrates with existing infrastructure.
In principle, direct-to-chip delivers fluid to cold plates mounted directly on the CPU/GPU, while immersion cooling submerges compatible hardware in a dielectric fluid that absorbs heat from the entire system. For business leaders, the meaningful difference isn’t the name of the technology but the investment architecture: direct-to-chip tends to be easier when retrofitting an existing data center, while immersion can be attractive when designing a new, very-high-density zone that needs to reduce dependence on airflow and optimize space. There is no single kW-per-rack threshold that applies to every project; the choice depends on GPU configuration, coolant temperature, redundancy requirements, maintainability, and the equipment ecosystem.
What’s notable is that cooling is starting to directly affect the economics of AI. A GPU only creates value when it is kept at high utilization and free from thermal throttling. Cooling has therefore stopped being an ancillary building item and has become part of “compute economics”: it affects how much GPU density can be deployed, floor space required, available power capacity, scalability, and ultimately the cost per unit of compute.
For Vietnam and other hot, humid markets, this issue deserves particular attention. When a business wants to bring AI to a private cloud or on-premise, it cannot simply buy GPU servers and drop them into an old server room. Power capacity, cooling, fluid distribution, floor loading, high-speed networking, storage, and monitoring all have to be assessed as one integrated system. This is exactly where immersion cooling can shift from being a “specialty” technology into a mainstream architectural choice for suitable high-density AI clusters.
Traditional enterprise infrastructure was built around the CPU — a general-purpose chip good at sequential tasks such as transactions, accounting, and administration. AI needs something different: massive numbers of computations running in parallel over huge blocks of data at the same time. GPUs do this far better than CPUs, which is why the GPU has become the heart of AI.
The key concept to remember here is the “memory wall.” For large AI models, what usually stalls the work isn’t compute speed but the capacity and speed of the GPU’s memory. The hardware progression below makes this clear:
| GPU generation | Memory | Memory bandwidth | Notes |
|---|---|---|---|
| H100 | 80 GB | 3.35 TB/s | Hopper architecture; first GPU with Confidential Computing |
| H200 | 141 GB | 4.8 TB/s | About 76% more memory and 43% more bandwidth than H100 |
| B200 (Blackwell) | 192 GB | ~8 TB/s | Draws roughly 1 kW per GPU |
Table 2. GPU memory — when a model “doesn’t fit,” this is the ladder you have to climb.

Figure 4. The “memory wall”: each GPU generation is, first and foremost, a leap in memory.
Beyond memory, the speed at which GPUs talk to each other decides everything. Within a single server, GPUs connect via NVLink at 1.8 TB/s; between different servers, they rely on InfiniBand networking at 400 or 800 Gb/s. Below these thresholds, distributed training across multiple machines slows down very quickly.
This shows that the GPU doesn’t create new problems — it exposes weaknesses that were already there in the infrastructure. A slightly slow storage system, a slightly congested network link — things an older system could “tolerate” — immediately turn into serious incidents once the GPU is running at full capacity. That is why so many organizations underestimate how hard it is to deploy AI.
Mr. Huang calls this layer the “AI factory.” A traditional data center handles business operations: email, accounting, storage, transactions. An AI factory manufactures intelligence — its output is prediction, inference, automation, optimization, and generated content. The difference between the two isn’t a minor upgrade; it’s a generational shift:
| Criterion | Typical data center | AI factory |
|---|---|---|
| Power per rack | 4–8 kW | 50–132 kW or more |
| Cooling method | Air | Liquid and immersion |
| Networking | 10–100 Gb Ethernet | 400–800 Gb + InfiniBand |
| Storage | Hybrid HDD/flash | All-flash, parallel file systems |
| Monitoring | Per-VM, focused on uptime | Per-GPU, per network port, per token |
| Power per facility | 1–2 MW | 10–50 MW or more per training cluster |
Table 3. The gap between a typical data center and an AI factory.
A common misconception is treating the network in an AI system as a single, uniform block. In reality, each training step passes through two distinct network layers: inside a single machine, GPUs connect via ultra-fast NVLink, with bandwidth that’s effectively “free”; between machines, an InfiniBand network carries the entire data exchange. It is the capacity of this inter-machine network layer — not the GPU itself — that sets the ceiling on large-scale training speed. That is why a single congested network port can stall an entire cluster running 1,000 GPUs.
Two real incidents illustrate this well. First, a 256-GPU cluster suddenly slowed by 35% overnight; engineers blamed the software, but the real culprit was a failing network connector that forced 255 healthy GPUs to wait on one weak link — the whole group can only run as fast as its slowest link. Second, expensive GPUs were running at only 40% utilization; the vendor advised “buy more GPUs,” but the actual cause was storage reading data too slowly. After moving hot data to local NVMe drives, those same old GPUs ran above 90% utilization — “buy more GPUs” turned out to be the most expensive way to fix a storage problem.
In practice, the infrastructure layer is rarely a pure choice between 100% “cloud” or 100% “on-premise” — it’s usually a mix of both: heavy training workloads go to major cloud providers (AWS, Azure, Google Cloud, Oracle, etc.), while inference work considered sensitive to data privacy is instead kept on-premise or on a private cloud. This hybrid setup always raises a fundamental security question known as the shared-responsibility model — who is responsible for what. The most dangerous gap usually isn’t inside the cloud or on-premise environment itself, but at the seam between the two.
For more than a decade, “cloud-first” was the prevailing principle of digital transformation: rent resources when you need them, scale up when demand grows. AI hasn’t devalued that model. On the contrary, public cloud remains especially well suited to fast experimentation, sharply fluctuating loads, short-term GPU needs, and organizations not yet ready to invest in dedicated infrastructure. But as AI moves from experimentation into production, the economics start to shift.
Inference workloads that run continuously create a cost structure different from experimental projects. Once demand becomes stable and GPU utilization is high enough, businesses can compare the cost of renting compute over time against the cost of owning the infrastructure. A 2025 Lenovo TCO analysis found that cloud has the advantage for short-lived or highly variable workloads, while on-premise infrastructure can have an economic edge for sustained GenAI tasks and continuous, high-utilization inference. This isn’t a universal rule for every business, but it shows that the “cloud vs. on-premise” decision is increasingly a workload-and-TCO calculation rather than an article of faith.
2026 market data also points to this rebalancing. Broadcom’s Private Cloud Outlook 2026 found that 56% of surveyed organizations are running or planning to run production AI inference on private cloud, while the equivalent public-cloud share fell from 56% to 41% year over year; 43% of organizations repatriating workloads said AI training, LLMs, and inference were among the workloads moved from public to private cloud. A separate survey published by Cloudian, covering 203 IT decision-makers, found that 79% had already moved part of their AI workload off public cloud or were in the process of doing so. These surveys don’t mean “cloud is ending”; they show cloud-first giving way to workload-first.
Five main drivers sit behind this shift. First, cost and budget predictability once inference becomes a continuously running background load. Second, data gravity: when enterprise data already lives inside factories, transaction systems, or internal data warehouses, continuously moving it to the cloud can add latency, bandwidth cost, and expense. Third, latency: manufacturing computer-vision use cases, real-time control, or edge AI need responses close to where the data is generated. Fourth, data sovereignty and compliance. Fifth, control: businesses want to know which model is running, where data goes, who has access, and the real cost per task.
However, “bringing AI on-premise” doesn’t mean moving everything back into a company’s existing server room. A genuine AI factory requires high-density power, liquid cooling or immersion cooling where appropriate, high-speed network fabric, sufficiently fast storage, orchestration, observability, and a specialized operations team. CIO.com has warned that repatriating AI into older data centers can carry significant retrofit costs. The realistic destination is therefore usually a spectrum of architectures spanning public cloud, private cloud, colocation, and on-premise AI factories.
A more useful framing for CEOs is not “cloud or on-premise?” but “which workload should run where?” Cloud can be the place for experimentation and absorbing demand spikes; private/on-premise can carry stable, sensitive, data-heavy, or latency-critical workloads; and hybrid becomes the orchestration layer between the two worlds. AI, then, isn’t reversing the cloud revolution — it’s making cloud strategy more mature.
This is exactly where the cloud-to-on-premise story meets the immersion-cooling story. When AI workloads return to infrastructure the business controls, the business also takes back responsibility for power, heat, networking, storage, and availability. In other words, greater control comes with greater infrastructure responsibility. A repatriation strategy that only prices in GPUs — without pricing in power and cooling — can easily produce an expensive data center that never realizes its full compute potential.
When people talk about AI infrastructure, the market often points to clusters of hundreds or thousands of GPUs, data centers rated at tens of MW, and large-scale liquid cooling systems. Those architectures are necessary for hyperscalers, sovereign AI programs, or organizations training very large models. But that is not the mandatory starting point for most SMEs and enterprises.
At business scale, many real use cases are far smaller: an internal AI assistant; RAG over a document repository; knowledge search; contract analysis; customer-support assistance; computer vision on a factory floor; security-log analysis; or an AI agent serving a specific process. NVIDIA is also introducing enterprise AI factory architectures designed to be modular and to scale unit by unit, identifying inference for small and mid-sized models, agentic AI, and industrial/physical AI as typical enterprise workloads. This shows that “AI factory” doesn’t have to mean hyperscale.
This tier of demand can be called Small AI: the business doesn’t start with “how many GPUs should we buy?” but with “which workload needs to run, what data needs protecting, what SLA needs to be met, and what level of performance is actually required?”
In the traditional deployment model, putting a GPU cluster inside a business can trigger a whole chain of investment: a server room, backup power, precision air conditioning, fire suppression, racks, networking, storage, monitoring, and an operations team. For a business that only needs a handful of GPUs or a modest AI cluster, that supporting infrastructure alone can become a major investment barrier.
A different approach is to package most of the AI infrastructure into a smaller deployable unit. A Small AI on-premise architecture can be pictured as four layers:
Immersion Cooling Tank → Server/GPU & Network → Private Cloud Platform → AI Models & Workloads
At the physical layer, immersion cooling can help manage high thermal density in a smaller footprint and reduce dependence on airflow in suitable configurations. However, immersion doesn’t remove every facility requirement: businesses still need to design proper power supply, heat rejection to the environment, electrical safety, fire suppression, floor loading, maintainability, and hardware compatibility. The Uptime Institute also notes that AI is driving adoption of both cold-plate and immersion cooling, but enterprise uptake still depends on thermal requirements, availability needs, and how well it integrates with existing infrastructure.
Above the physical layer sit the servers, switches, and a private-cloud platform. With platforms such as VMware, Virtuozzo Hybrid Infrastructure, or OpenStack, vendor documentation confirms support for PCI/GPU passthrough and NVIDIA vGPU when compute nodes are configured appropriately. GPU passthrough lets a virtual machine access a physical GPU directly assigned to it; vGPU allows a GPU to be shared according to profiles supported by the hardware and software.
The important architectural point is that a business doesn’t have to standardize an entire cluster on the single largest, most expensive GPU. In principle, compute nodes can carry different types of GPUs, with a scheduler/private-cloud platform assigning workloads to the node with suitable resources. That said, exact support has to be validated for each specific GPU, server, driver, firmware, hypervisor, and licensing mechanism; “heterogeneous GPU” should not be read as meaning any GPU type can be freely swapped within the same distributed workload.
One of the biggest sources of AI waste is over-provisioning: using a premium GPU for a workload that doesn’t need that level of capability. Inference for a small model, internal RAG, computer vision, embeddings, fine-tuning, and large-scale training all have very different requirements for VRAM, memory bandwidth, throughput, and latency.
Small AI should therefore follow this decision chain:
Workload → Model → Right-sized GPU → Right-sized infrastructure
instead of:
Buy the biggest GPU → then look for a workload to run on it
A light workload can run on a smaller GPU; a workload that needs more VRAM is assigned to a GPU with larger memory; a workload that needs high throughput or low latency is reserved for a more powerful GPU. When private cloud, GPU passthrough/vGPU, and orchestration are designed correctly, a business can build a multi-tier pool of AI resources instead of forcing every application onto the same expensive configuration.
For leadership, the KPI should therefore shift from “which GPU is fastest?” to “what is the lowest cost to complete this workload within the required SLA?”
The push to bring AI on-premise in Vietnam isn’t driven by cost alone. The Digital Data Law No. 60/2024/QH15 has been in effect since 1 July 2025, and the Personal Data Protection Law No. 91/2025/QH15 has been in effect since 1 January 2026, making data governance, protection, and lifecycle control an increasingly important issue for businesses.
This needs to be stated carefully: Vietnamese law should not be read simplistically as “every piece of data from every business must be stored inside Vietnam.” Domestic-storage obligations depend on the type of data, the entity involved, and the specific case. For example, Decree 53/2022/ND-CP sets out domestic data-storage obligations for certain categories of businesses and data falling within the scope of the Cybersecurity Law, subject to specific conditions and procedures.
Still, even when a workload isn’t legally required to be fully on-premise, many businesses have good governance reasons to keep important data and AI inference within infrastructure they directly control: fewer third parties touching the data, simpler access-control requirements, direct ownership of audit logs, and compliance with internal data-residency policies.
At the same time, on-premise brings compute closer to the data. For factory-floor computer vision, AI-assisted operations, RAG over large document repositories, or agents accessing internal systems, processing near the data source can reduce data movement, reduce reliance on internet connectivity, and improve latency.
This opens up a notable architectural middle ground between two extremes — “renting all of your AI on public cloud” and “building a large-scale AI data center”: a small-to-medium private AI infrastructure, integrating immersion cooling, servers/GPUs, networking, and a private cloud, sized to actual workload needs.
It can be pictured as an “AI Box” or “Micro AI Factory” installed at the business. When demand is still small, the organization starts with one or a few GPU nodes. As the number of models, users, or workloads grows, the infrastructure can scale by adding nodes; public cloud continues to be used for experimentation, burst capacity, or specialty GPUs not yet worth investing in on-site.
The target architecture, then, doesn’t have to be cloud or on-premise — it can be:
Small AI on-premise + private cloud + public cloud when needed.
Lenovo’s 2025 TCO analysis found that cloud still has the advantage for short-lived or highly variable workloads, while on-premise infrastructure can achieve better cost efficiency when AI workloads run continuously and GPU utilization stays high. The study itself also notes that real-world TCO still depends on networking, facilities, operations staffing, and many costs beyond compute alone. So the goal isn’t to prove that on-premise is always cheaper than cloud, but to find the point at which utilization is stable enough for infrastructure ownership to start making economic sense.
Market trends are also moving in a workload-first direction. Broadcom’s Private Cloud Outlook 2026, surveying 1,800 IT leaders across eight countries, found 56% of organizations running or planning to run production AI inference on private cloud, while the public-cloud share for the same workload fell from 56% to 41% year over year. A 2026 Cloudian survey of 203 IT decision-makers likewise found that 79% had already moved part of their AI workload off public cloud or were in the process of doing so. This is vendor survey data, not proof that every business will leave the cloud; but it reinforces the point that cost, control, data sovereignty, and performance are pushing AI architecture toward greater distribution.
If hyperscale AI is a story about thousands of GPUs and tens of MW of power, Small AI is a story about bringing dedicated AI capability closer to the ordinary business. The goal isn’t to build the largest possible infrastructure, but infrastructure that is just right — secure, governable, and able to scale.
For CEOs, the key message is this: on-premise AI doesn’t have to start with a multimillion-dollar data-center project. A business can start small, keep its most important workloads and data within its own control, choose GPUs to match the model and workload, scale as demand grows, and keep using public cloud wherever cloud has the advantage.
Public attention still revolves around chatbots, but how businesses actually deploy AI is a different story. The largest long-term value is likely to come from operational and physical AI: factory automation, predictive maintenance, production optimization, digital twins of equipment, cybersecurity automation, medical analytics, and infrastructure operations intelligence. Mr. Huang has also emphasized the opportunity in robotics and “physical AI,” where manufacturing capability is combined with AI.
These operational AI applications are far more demanding than a chatbot. Factory AI needs absolutely consistent low latency; medical AI needs traceability and auditability; cybersecurity AI needs real-time data; infrastructure AI needs a continuous stream of monitoring. In other words, the quality of a model depends directly on the quality of the infrastructure feeding it.
| The model itself is not a durable competitive advantage; increasingly, it is the operating environment that nurtures the model that becomes the real advantage. Once foundation models are “good enough to build on,” the difference lies in who operates them better. |
The application is the most visible layer, because it’s where a business directly feels the results: virtual assistants, process automation, predictive analytics, intelligent request routing, automated incident correlation, infrastructure-optimization tools, and more. Mr. Huang himself has affirmed that this top layer is ultimately where the real economic benefit shows up.
But an application only succeeds when the layers beneath it are mature enough. This is where many AI projects fail: a business rushes to deploy an application while data is still fragmented, monitoring is still rudimentary, infrastructure bottlenecks haven’t been identified, and governance mechanisms aren’t yet operating. The outcome is almost predictable: unreliable answers, inconsistent performance, users losing trust, security incidents, and infrastructure costs spiraling out of control.
The organizations that see clear returns take a different approach: they treat AI first as an infrastructure-modernization project, and only then as an application project.
This is the part many write-ups leave out, and the point that most needs emphasis: security isn’t a “sixth layer” stacked on top; it’s a vertical thread present in every layer — because an attacker can target any of the five layers.
| Layer | Characteristic attack type | Reference framework |
|---|---|---|
| 1 — Power | Physical sabotage, cutting power or cooling, malicious hardware planted in the supply chain | Data-center physical security |
| 2 — Chip | Data theft via GPU memory, malicious firmware | Confidential Computing |
| 3 — Infrastructure | Lateral movement, cloud misconfiguration, shared-responsibility gaps | Zero Trust |
| 4 — Model | Training-data poisoning, model evasion, model theft | MITRE ATLAS, NIST |
| 5 — Application | Prompt injection, sensitive-data leakage, excessive AI privileges | OWASP Top 10 for LLMs |
Table 4. Risk map across the five layers, compiled from multiple sources.

Figure 5. Security is a vertical thread running through all five layers — each layer has its own risk type and defense framework.
The severity of these risks is rising fast. MITRE ATLAS — a database of AI attack techniques, similar to a familiar cybersecurity playbook — had, as of November 2025, catalogued 16 tactic categories and 84 techniques, with an early-2026 update adding techniques targeting autonomous, self-acting AI. NIST groups attacks on AI into four main categories: evasion, poisoning, privacy compromise, and abuse.
Previously, data could only be protected in transit and at rest; data being actively processed remained exposed. This is a critical weakness for AI, because processing is exactly when sensitive data (personal information, trade secrets) and high-value models are most exposed.
The NVIDIA H100 was the world’s first GPU to address this through Confidential Computing. The mechanism: it creates a “secure enclave” directly in hardware, isolating the entire workload; combined with secure boot (only verified software can run) and hardware attestation (cryptographic proof that the GPU is in a secure state). To be effective, the server also needs a CPU that supports the corresponding secure-enclave capability from AMD or Intel.
The biggest concern when enabling this feature is “we get security, but how much slower does it get?” The research findings are encouraging: for large-language-model inference on H100, the performance overhead is typically under 5%, and close to zero for large models and long sequences; most of the slowdown comes from encrypting data as it moves between CPU and GPU, not from the computation itself.
From a compliance standpoint, Confidential Computing fills exactly the “data in use” gap that healthcare and financial-data regulations require — allowing inference to run on shared infrastructure without exposing data to the cloud provider, administrators, or an attacker. It’s worth noting this is a technical component, not a standalone certification; it must be paired with encryption, access control, and audit logging to form a fully compliant inference stack.
The hybrid infrastructure layer demands a “trust no one by default” mindset (Zero Trust): never trust a connection just because it sits inside the internal network; always authenticate and grant only the minimum privileges needed for every data flow. This matters even more as AI drives a surge in east-west traffic between GPUs — the same network path that carries performance can also become the path an attacker uses to move laterally.
In the cloud, responsibility is split down the middle: the provider secures the cloud itself (hardware, facilities), while the customer secures what’s on the cloud (configuration, data, identity, models). With AI, there’s an added question: who controls the model and the data during inference? Confidential Computing, discussed above, is precisely the tool that pulls this trust boundary back toward the customer, even when running on shared infrastructure.
At the top of the stack, the most widely used reference framework is the OWASP Top 10 for LLM Applications, 2025 edition. A few highlights:
Excessive agency: as AI shifts from “just answering” to “acting autonomously” (sending emails, querying databases, calling services), granting it more privilege than necessary creates a vulnerability.
Prompt injection has ranked #1 for two editions running. The root cause is that a model reads instructions and data through the same “channel” without a clear boundary between them — so an attacker can craft text that tricks the model into treating data as a command. The most dangerous variant is indirect injection, hidden inside a document that the AI reads.
Two new risks in 2025: system-prompt leakage, and weaknesses in vector-database stores — reflecting how widespread retrieval-augmented generation (RAG) architectures have become.
Excessive agency: as AI shifts from “just answering” to “acting autonomously” (sending emails, querying databases, calling services), granting it more privilege than necessary creates a vulnerability.
These are not theoretical risks. Even leading models can still be successfully compromised through indirect prompt injection, and the success rate rises with repeated attempts — showing that single-layer defense isn’t enough. OWASP is explicit: no single solution eliminates this class of attack entirely; defenses must be layered — least-privilege access, filtering both inputs and outputs, human approval for high-risk actions, and regular adversarial testing. This practice is giving rise to a new discipline, sometimes called MLSecOps — embedding security across the entire AI lifecycle.
The OWASP list identifies technical risks; NIST’s AI Risk Management Framework (AI RMF) and the ISO/IEC 42001 standard provide a way to manage them systematically and auditably — each OWASP risk can be mapped to controls within these frameworks. A complete governance architecture typically combines three pieces: a map of attack techniques (MITRE ATLAS), a prioritized risk list (OWASP), and an auditable governance framework (NIST, ISO 42001).
For many organizations, this governance layer also has to align with data-protection and data-sovereignty regulations — requirements that citizen, patient, or customer data be processed and stored within specific legal borders. This regulatory trend (from GDPR in Europe, personal-data-protection laws across Asia, to the EU AI Act) converges with Mr. Huang’s view that every country should build its own AI capability rather than simply buying it. Technically, the combination of Confidential Computing, on-premise or private-cloud deployment, and a Zero Trust mindset is the toolset that helps reconcile three goals that often conflict: GPU performance, data sovereignty, and legal compliance.
AI infrastructure poses a “visibility” challenge that traditional monitoring isn’t built for. Conventional monitoring revolves around uptime, CPU usage, and storage capacity. AI environments need to go deeper: measuring the latency of every response, how busy each GPU is, network congestion, token-generation speed, and detecting when a model has “drifted” from its original accuracy.
A key security point: many signals are simultaneously operational and security-relevant. A sudden drop in accuracy could be natural data drift — or a sign of data poisoning. Unusual network traffic could be a technical bottleneck — or an attacker moving laterally. So modern monitoring tools must serve both operations and security teams, and must sample at sub-second granularity, since many phenomena last only milliseconds before disappearing. The overall trend is shifting from “monitor to react” toward “predict to prevent.”
When the five layers are combined with the security thread running through them, AI becomes a question of capital allocation, operational capability, and risk governance — not just a technology choice. Before committing to a large-scale AI program, leadership should demand clear answers to five questions: