AGIX Portfolio Manager Notes

The Agent Handover: Computer Use, Sandboxes, & The Emerging CPU Opportunity

By Max Chen & Derek Yan, CFA

In our last KraneShares Public-Private AI & Technology ETF (AGIX) Portfolio Manager Note, we argued that being "realized" is the central storyline of the intent agent, the agent that can continuously advance a user's goals on their behalf. This time, we turn to the interface through which an intent agent does its work and to the opportunities that interface may create.

Agents Could Be Using Computers Like People

Steve Jobs liked to call the computer a bicycle for the mind.1 People act on the world in two ways: by doing physical work, and by producing knowledge in the digital realm. For decades, the primary interface to that knowledge world has been the personal computer. We believe AI agents may retrace a similar evolution, from today's command-line era to a graphical one.

This is not to say that code cannot get the job done. Generating a PowerPoint, editing an Excel chart, even directing a robot: in each case, code is already serving as the universal operating language. But there are two reasons agents need to operate graphical interfaces as well. First, not all existing software was built around code or around agents and APIs (application programming interfaces, the structured channels through which programs communicate). When a task spans several applications and a long time horizon, the graphical interface becomes an API that requires no vendor's permission. Second, however reliable agents become, we think people will still need visibility into the work as it unfolds. Visibility means being able to step in and take over at any moment, which, in turn, requires that the agent and the human operate the same interface. (Agents may ultimately keep two: code and the screen.)

This is why the frontier labs are pushing so hard on computer use. OpenAI's newest model, Astra, achieved higher computer-use performance in OSWorld (a benchmark that tests whether an AI agent can operate web and desktop applications in a real computer environment and complete tasks) while taking roughly 47 percent less time per task than GPT-5.6 Sol.2

Progress in computer use could come from the supply chain that gives models a scalable interaction experience, which is to say, data. We think short-task benchmarks already show strong capability, but completion rates on long, end-to-end workflows remain low, and that gap is generating demand for real workflows, reinforcement-learning environments, and reliable verifiers (the components that judge whether a task was actually completed and turn that judgment into a training signal). Several public companies are already aiming at personalization for long-horizon agents and at building reinforcement-learning environments for computer-use agents.

Sandboxes May Be the Main Infrastructure Opportunity for the Growth of Computer-Use Agents

In the training phase, when a reinforcement-learning environment (RLE) is being assembled, infrastructure providers such as California-based Sandbox Services come into play. A sandbox is a service that gives an AI agent an isolated computing environment: for each training attempt, it launches a working browser or virtual desktop in which the model can observe the screen, click, type, and carry out a task; when the task ends, the environment is reset to a preset state so the experiment can be repeated. A training system also needs to run vast numbers of these environments in parallel, collect the resulting action trajectories, and have a task verifier score the outcome and generate the reward signal. Startups such as Daytona, E2B, and Modal provide these services, and their customers are, for the most part, other startups whose main business is selling RLE data to model companies.

We think training has to be resettable, and deployment has to be persistent. Once an agent moves into real-world deployment, the picture changes. Meta's own account of Muse (its personal agent app), each user gets a dedicated virtual machine in the cloud: the agent uses a browser and tools inside it, and handles the user's files there.3 At the same time, Meta* (a public holding of AGIX) isolates the agent's runtime from credentials and security services, and routes every external action and network request through a separate permission-control component called Sentinel. Reliability, state management, security, and permission governance could, in principle, all be solved through the technical progress of third-party sandbox providers. But we believe that when agents are deployed to consumers at scale, cost and the boundaries of liability may outweigh purely technical considerations, potentially leading companies like Meta, Perplexity, and SpaceXAI* (also an AGIX holding) to build this infrastructure themselves to meet their products' needs.

Perplexity reported that SPACE created sandboxes three to five times faster than its previous sandbox provider on the same production traffic.4 This tension between buying and building is, in our view, the single largest risk in investing in these infrastructure startups. The agent runtime may not be an infrastructure problem at all, but a product problem. And if it is a product problem, it should ultimately belong to the product companies, not to an independent middle layer.

At the same time, E2B and Modal rely on infrastructure from major cloud providers: E2B's managed sandboxes run on Google Cloud, while Modal uses infrastructure from AWS and Oracle*, among others. That reliance creates a potential competitive vulnerability, as the underlying providers can also offer native sandbox services. AWS already does this through Amazon Bedrock AgentCore Code Interpreter.

In our view, hyperscalers could price these services aggressively to drive demand for their broader cloud infrastructure, rather than prioritize standalone sandbox profitability. Independent providers would then face competition from the same companies supplying their underlying compute. The investment opportunity may therefore extend beyond the sandbox itself to the CPU-hours consumed as agents execute tasks. For that reason, we remain more constructive on the growth potential of incumbent cloud giants like Google*, Amazon*, and Microsoft* (which AI ETF AGIX currently overweights as of September 30, 2026), and on the additional CPU demand that broader adoption of computer-use agents could generate.

Why We Believe The CPU Matters For Computer-Use Agents

In Perplexity's post on SPACE, its sandbox platform, the company noted that in early tests on the NVIDIA Vera CPU, real workflows ran roughly 1.5 times faster than its current production reference, and concurrent sandboxes started up to 1.9 times faster.4 Vera is NVIDIA's Arm-based CPU, and it differs sharply from the traditional x86 processors (the architecture from Intel and AMD that has dominated servers for decades). NVIDIA*, a public holding of AGIX, offers an apt description of its own: an agent workload is a long serial chain of reasoning, punctuated by sporadic bursts of parallel work.5 The serial main path is strictly latency-bound and determines how long the entire session takes to complete.

We believe these workloads place two potentially competing demands on a CPU: strong single-thread performance to handle the critical path and enough cores to support large numbers of concurrent microVMs (lightweight virtual machines), ideally with predictable latency as utilization rises. Traditional x86 designs often combine high core counts, multiple threads per physical core, and distinct local memory regions within a processor. In our view, these features can deliver strong performance at lower utilization, but as concurrent tasks put pressure on shared cores, caches, or memory channels, competition for resources may increase, potentially extending completion times for the slowest requests.

We think Arm-based CPUs like NVIDIA's Vera, which use architecture from British semiconductor and software design company Arm*, could offer an alternative approach to some of these challenges. NVIDIA argues that Vera is particularly well suited to agent workloads; in our view, its potential advantages will depend on how it performs under real-world concurrency and utilization conditions.

Meanwhile, as agents of every kind proliferate, Intel's chief executive, Lip-Bu Tan, has said that the ratio of CPUs to GPUs has moved from 1:8 to 1:4 and could move toward 1:1.6 Arm, a public holding of AGIX, expects data centers to require more than four times today's CPU capacity per gigawatt as agentic AI scales, creating what it estimates is a market opportunity of more than a 100 billion dollars by 2030.7 Each of these projections elevates the CPU to a position of unusual importance. We believe that, as OpenAI, Anthropic* (private holding in AGIX), and other model companies may follow with always-on agent products of this kind, the market may reprice CPU-related companies.

The Opportunities & Challenges Of Migrating From Computer-Use To Phone-Use Agents

The more we have used Muse, the more we have noticed its limits. To begin with, the evolution of the user interface from the command line to the PC is only an intermediate stage; mobile-native behavior is already embedded in the daily habits of ordinary users. When the browser take-over experience it offered on the phone was noticeably imperfect, users struggled a little bit as if they needed to operate a PC on their phone. Most personal services, from ride-hailing and grocery delivery to booking hotels and flights, already run smoothly through apps, and some exist only as apps, with no web interface for an agent to operate.

Does that imply that a boom in phone-use agents may well follow the boom in computer use?

The phone-use model could look very different from the computer-use model. Muse's architecture works because web endpoints are open: anyone can spin up a temporary browser in the cloud, and once logged in, the agent is, for all practical purposes, the user. Mobile apps' risk-control systems, by contrast, treat the device itself as part of a user's identity, and that part cannot be replicated in the cloud. ByteDance's Doubao phone (Doubao is ByteDance's consumer AI assistant; the phone was built by renting the system layer of a second-tier handset maker) took this route earlier. But every app outside ByteDance's own ecosystem fought the phone-use agent, and the product launch faded without result, even for a player like ByteDance, with its potential to build an agent operating system on top of the human operating system.

Our call is that phone-use sandboxes still have a market, but one that is likely to remain centered on model training, along the lines of the Android virtual machines run through QEMU (an open-source emulator) that DeepSeek describes in its DSec paper, its account of the sandbox platform behind its agent training.8 In deployment, the path most likely runs through the operating system: exposing structured, operable app intents, as AGIX public holding Apple* does (a framework through which apps declare specific actions the system can invoke on their behalf), and using small models that can run on the device (in the spirit of Jev, TypeSafe AI's non-generative decision model released last week, which scores candidate actions rather than writing text) to adjudicate and backstop each step.

And where a task must cross applications and runs into entities such as WeChat or Amazon, which would prefer to keep their relationship graphs and product catalogs inside their own walls, their products and services may instead be reached agent-to-agent, preserving a seamless "L4" agent experience (borrowing the autonomous-driving scale, where Level 4 means the system completes the task within a defined domain without human intervention).

With that in mind, while significant roadblocks remain and numerous technical challenges must be solved to serve billions of people with always-on agents, we believe we are already witnessing new opportunities emerge. The agent can act on our behalf, and the handover could have profound implications for both established and emerging systems, which we will aim to capture such opportunities in AGIX. 


*AGIX holdings are as of September 30, 2026, and are subject to change. Securities discussed do not represent the Fund's entire portfolio.

For AGIX standard performance, top 10 holdings, risks, and other fund information, please click here.

Citations:

  1. The Marginalian, "Steve Jobs on Why Computers Are Like a Bicycle for the Mind (1990)," retrieved on 9/30/2026.
  2. OSWorld, "OSWorld 2.0: Benchmarking Computer-Use Agents on Long-Horizon Real-World Tasks," retrieved on 9/30/2026.
  3. Meta, "How We Built Safety Into Muse," as of 9/8/2026.
  4. Perplexity, "Secure sandboxes for agents," as of 7/15/2026.
  5. NVIDIA Developer, "Solving Agentic AI Fleet Challenges with NVIDIA Vera CPU," as of 8/24/2026.
  6. Investing.com, "Earnings call transcript: Intel's Q1 2026 earnings beat forecasts, stock rises," as of 4/23/2026.
  7. SEC Form 6-K for Arm Holdings plc, as of May 2026.
  8. Arxiv, "DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale," as of 9/19/2026.

Definitions:

Command-Line Interface (CLI): A text-based computer interface in which a user or program runs commands by typing them into a terminal rather than interacting with visual elements such as windows, icons, and menus.

Graphical User Interface (GUI): A visual interface that allows users to operate software through on-screen elements such as windows, menus, buttons, icons, and text fields.

Intent Agent: An AI agent designed to pursue a user's ongoing objective with a degree of autonomy, potentially selecting and executing actions over time rather than responding only to one-off prompts.

Long-Horizon Agent: An AI agent intended to complete tasks that require many sequential actions, extended time periods, or coordination across multiple applications and services.

Multimodal Model: An AI model that can process or generate more than one type of information, such as text, images, audio, video, or computer-screen inputs.

Reinforcement Learning (RL): A machine-learning method in which an agent improves through trial and error by receiving rewards or penalties for its actions.

Reinforcement-Learning Environment (RLE): A real or simulated setting that defines what an agent can observe and do, how its actions change the environment, and how success is scored.

CPU-Hour: One CPU core running for one hour, commonly used as a measure of cloud-computing consumption and billing.

MicroVM: A lightweight virtual machine that provides stronger isolation than a container by running its own kernel while aiming for relatively fast startup and low overhead.northflank+1

Node: A single server or computing instance within a larger cluster of connected machines.

Virtual Machine (VM): A software-defined computer that emulates a separate operating environment, including an operating system and allocated compute, memory, storage, and network resources.

Identity and Access Management (IAM): The policies, systems, and controls used to determine who or what can access a resource and under which conditions.

L4 Agent: An analogy to Level 4 autonomous driving, describing an agent that can complete a task without human intervention within a defined operating domain and set of conditions.