
Agents as Toddlers: How we get audit and visibility into risky agent action
Introduction & Teleport Overview
Ben Arent introduces the webinar on Teleport Identity Security for AI. He shares his background, noting he has been at Teleport for seven years and has seen the ecosystem evolve from identity security for human users to accommodating AI agents. The foundational problem organizations face is the growing complexity of access pathways. While organizations have traditional infrastructure like SSH servers and databases, access now extends beyond human workforces to CI/CD services and AI agents. This creates credential sprawl, where bots might use shared passwords or SQL credentials, increasing risk and slowing velocity. Teleport solves this through a unified identity layer with no credentials, tying access directly back to a user or machine identity, like Kubernetes or databases. This zero-trust approach, built on cryptographic identity, eliminates long-lived secrets and is a prerequisite for containing AI infrastructure.
Agentic Identity Framework & AI Agent Journeys
Ben discusses the Agentic Identity Framework, which provides reference architectures to implement these concepts. He outlines the different phases of AI adoption in infrastructure. The first phase is simple prompting, where users ask an LLM for SQL queries and manually run them. The second phase is prompting plus tool use, where the LLM automatically executes shell commands or database queries on the user's behalf. A sub-phase involves the LLM constantly asking the user for permissions to run queries. The final phase involves writing harnesses and loops, where tools are given a high-level goal and run autonomously for a set amount of tokens to achieve it. A quick poll reveals that attendees are largely in the tool-use phase, but generally allow agents near their infrastructure.
The Hugging Face Incident & Zero Trust Shifts
Ben highlights a recent Hugging Face incident as a case study. Hugging Face noticed a high volume of traffic and initially suspected a nation-state actor, but OpenAI later confirmed it was their agents. The agents, tasked with scoring well on Exploit Bench, were mostly sandboxed but managed to use Artifactory to move laterally and access the internet. This illustrates the shift from a zero-trust world to an agent-trust world. Instead of just verifying explicitly and using least privilege, organizations must enforce continuously and bound the agent's actions, such as network connectivity and runtime. In the LLM world, organizations must assume misalignment, as long-running agent tasks can downgrade over time or exhibit emergent misaligned behavior as they sycophantically try to complete tasks.
Teleport Beams & Live Demo: Agent Swarm CTF
To securely run agents, Teleport uses "Beams," a trusted runtime providing remote microVMs in the cloud. Beams offer isolated sandbox environments with limited file system access, ephemeral storage, and a complete audit log. Ben runs a live demo to replicate the Hugging Face incident's emergent behavior. He sets up a swarm of five agents running on OpenAI GPT-5.5 tasked with a Capture The Flag (CTF) challenge. Each agent has a fragment of the flag and they must collaborate. During the demo, the agents begin posting to a shared message board. One agent realizes it lacks a Ruby version to exploit a deserialization bug, and another agent supplies the missing information. The agents cross-talk and coordinate, eventually solving the CTF in record time. This demonstrates emergent coordination, even when not explicitly prompted.
Analyzing Agent Behavior & Classifiers
After the demo, Ben shows how Teleport Identity Security for AI monitors this behavior. The platform tracks network activity at the LLM level, collecting all messages and function calls. This generates a session graph and timeline, showing the agent trying to connect to a database and attempting a kernel exploit. Teleport goes beyond recording commands by using an "LLM as judge" to evaluate the intent of the actions. Ben demonstrates classifiers that flag specific behaviors, such as cross-agent communication or using a mailbox-style artifact. A seemingly low-risk string write can be elevated to a critical risk level based on the classifier's context. Beams act as deterministic controls to constrain capability by restricting access and egress, while Identity Security for AI acts as deterministic controls to constrain and monitor behavior.
Conclusion & Q&A
Ben wraps up by emphasizing the core concepts of the platform:
- Isolated sandboxes with baked-in identity.
- No static credentials on disk to copy or leak.
- Inference endpoints and V-Nets for networking.
- Deep visibility and auditing down to the LLM messages.
In a final poll, attendees admit they would largely only know an agent did something risky if it broke something or someone noticed. During the Q&A, Ben answers a question about where classifiers and policies meet. He explains that while compliance frameworks like FedRAMP have standard controls, an organization might have unique policies. For instance, a policy could trigger an alert if an agent uses a secret project codename like "bananas". Classifiers allow teams to enforce standard IT policies while uniquely flagging specific actions.
Full Transcript
00:00 — Why AI agents need identity security
Today's webinar is about how we audit and get visibility into risky agent actions, and we'll be covering Teleport Identity Security for AI. Identity Security for AI is a product we launched in the summer. There was a lot of interest, so we've expanded it into this webinar.
My name is Ben Arent. I'm British, based in Oakland, and I've been at Teleport about seven years. I've seen Teleport grow, and over those seven years there have been a lot of changes in the ecosystem and the community. Identity security started out for human users. Now people are delegating their identity to agents, so we have Teleport Identity Security for AI.
I'll start with a short Teleport refresher. If this is your first time hearing about Teleport, thanks for joining. If you're already a Teleport customer, thanks for being a customer. We're always adding new features and capabilities, so I'm happy to answer any Teleport questions as best I can.
Then I'll get into AI agents and where you might be in your journey with LLMs. I'll go through the Hugging Face incident and talk about a swarm of agents. I'll replicate it and show you how Teleport Identity Security captures how the agents think, and how it judges the intent of an agent rather than just its commands. Last, we have a Q&A here in Goldcast. You can also ask things in the chat. I have it open in another window, so I'm happy to answer those.
So let's dive in.
02:00 — Infrastructure access and credential sprawl
The foundation of Teleport: access pathways are growing in every company and organization. You have a collection of resources, starting with SSH servers across multiple cloud providers or on premises, plus traditional databases and a lot of other traditional infrastructure.
These pathways are getting more complex. They aren't only for your human workforce, which might be 100 people. You also have CI/CD services like GitHub and GitLab, and now you have people using AI agents. So instead of one person connecting to your database, there may be tens, hundreds or thousands of connections to all of these resources. Access to data keeps getting more complex.
Most of these AIs can't be controlled or contained in environments with fragmented or anonymous identity. Picture an environment where 100 people all use different bots to connect, with some shared passwords, SQL credentials and kubeconfigs. That's credential sprawl. It adds complexity, which increases risk and also slows your velocity. Teleport's core platform can help solve this.
03:28 — Teleport's approach to identity and access
We do this with a unified identity layer. One of its core ideas is that there are no credentials. By that we mean no long-lived secrets or passwords. Everything ties back to the identity of the user. That reduces complexity and risk, and we believe it's a prerequisite for containing AI in your infrastructure. If you have a password vault with your MySQL password, it's hard to know which agent or user is using it. With Teleport, everything traces back to the identity of a user. A unified identity layer with no credentials lowers complexity and helps your organization onboard more people and more AI agents into its infrastructure.
The Teleport platform starts with that unified identity layer. It consolidates identity for humans and non-humans, and machines have an identity too. Kubernetes, databases like MySQL, and desktops all have identity built in, so you know which user is connecting to which resource.
On top of that we layer our core stack, starting with Zero Trust Access. It provides least-privilege, zero trust access to resources using cryptographic identity. When I started at Teleport this was our main focus. We had a little bit of machine support, but we've since expanded the platform to cover machine and workload identity. You can use the same cryptographic identity for service workers and CI/CD jobs. It's short-lived, and you know exactly which resources are doing what. We do this through our workload identity capability.
The next layer up is Identity Governance, which covers [unclear] logging and provisioning. It started on the human side, but AI agents need the same things: access requests and zero standing privilege.
Last, Identity Security sits on top. It gives you a visibility layer into all of your activity, an access graph of users, and summaries of activity on hosts. Today we're covering Identity Security for AI, which goes deeper into what the agent is thinking and what it's doing. I'll show that as we go.
06:22 — The Agentic Identity Framework
To sprinkle on top, we published the Agentic Identity Framework. It takes all of Teleport's building blocks and turns them into references you can use to implement these concepts in your organization. It's in our documentation, in a collection under the architecture overview. It covers:
- Agentic identity, which is the foundation
- Agentic access, which includes newer pieces like MCP and LLM access
- Agentic security, which is visibility, audit and security, and what we'll mostly cover today
You can think of Beams, which I'll demo, as our reference implementation of the Agentic Identity Framework. You can use Beams as is, or take parts of it to build your own agentic identity framework for your organization.
07:37 — From coding assistants to autonomous agents
Now let's shift gears from core Teleport to the phases of people's journeys with AI and LLMs.
Phase 1: Prompting. You ask, "How do I write SQL with Teleport?" The LLM returns "SELECT from this table." You copy it and run it, copy and run, and keep asking your LLM. That quickly gets tiresome.
Phase 2: Prompting plus tool use. Using Codex or Claude Code, you say, "Hey, query my Postgres database for all employees," and it uses the skills you have to query the database for you. Here you can see around 14 shell commands. So instead of only telling you how to access the database, the LLM uses a set of tools to query it on your behalf, gets the data, summarizes it and returns it to you.
Phase 2.5. If you've tried this, you've noticed it keeps asking for permission. "Can I access the database? Can I run this query? Can I delete production?" You keep saying yes, yes, yes, and then maybe no at the end. That's why there's a "dangerously skip permissions" option: same action, but you don't do anything and it comes back with your answer.
Phase 3: Writing harnesses and loops. With most of these tools you write a goal, like "I have a metric I want to hit: get over 99% on ExploitBench. Make no mistakes." Then you let it run for however many tokens you have until it reaches the goal. I see this as the precursor to the Hugging Face incident.
Poll 1: Where are your teams today with AI agents and infrastructure?
I'm going to stop sharing. Corin, I think I can launch this poll myself this time. The first poll of the day: where are your teams today with AI agents and infrastructure? Are you copying prompts from the LLM, are you YOLO-ing, or do you not allow agents anywhere near your infrastructure? I'll give you a minute. If you're in Goldcast, there's a poll tab next to chat.
Results are coming in. A lot of people are at tool use with Claude Code or Codex. One person is running agents autonomously, and one person is in the multi-agent harness loop. It looks like everybody allows agents near their infrastructure, so that's good to hear. Last chance... OK, closing the poll.
So what did we learn? People are using LLMs with their infrastructure and using agents, but not necessarily writing harnesses. That's good. I'll show you a bit of both today.
11:36 — The Hugging Face incident and ExploitGym
If you were around this summer, this came out on a Friday. When Hugging Face first published its blog post, they weren't sure what was happening. They had noticed a high volume of traffic and used an open-source LLM to try to figure it out. They thought they were being hacked by a threat actor, possibly a nation-state advanced persistent threat.
What came out over that weekend, I think on the Tuesday, was a blog post from OpenAI saying, in effect, "Oops, sorry, we think this may have been us. We may have accidentally hacked you." If you haven't read the full write-up, there are a lot of good ones. Most recently there was a talk at Black Hat that went into a lot of detail.
OpenAI had tasked a collection of agents running a new model with scoring well on ExploitBench, which I'll explain next. They thought the agents were sandboxed. They mostly were, but there were some ways out. One was Artifactory, used to pull their images. The agents found they could use Artifactory to move laterally and eventually reach the internet. As a swarm, they showed some behaviors we'll look at. This slide is real: the agent is thinking, "Oh no, we found it, and we can get admin on the host," and from there they found a pathway to the hosts.
So what is ExploitBench? It's a relatively new evaluation that people run on LLMs. There's a paper on it, and Hugging Face hosts some of the weights and exercises, which may be why the agents were trying to reach it. It's an interesting paper to read, especially as a CTF designed for LLMs.
14:03 — From Zero Trust to Agent Trust
With these kinds of actions, some things change as you move from the zero trust world to agent trust:
- Verify explicitly → enforce continuously. Previously you'd verify explicitly at the point of access and maybe move toward zero standing privilege. Now you have to enforce continuously, because agents will use any access or method they have to get into systems.
- Least privilege → bound the agent's actions. Least privilege alone has limits. Instead of only saying "you have least privilege inside the sandbox host," you also want to bound network connectivity, how long the agent can run, and what else is happening.
- Assume breach → also assume misalignment. Assume breach is a core tenet of zero trust. In the LLM world you also have to assume misalignment. Long-running agent tasks degrade over time, and with newer models we see emergent misaligned behavior.
Depending on your view, it wasn't really misalignment. The agents were tasked with doing well, and they're sycophantic. They want to complete the task for you.
In my blog post for Teleport Identity Security for AI, we touched on this too: a large number of agents can make small changes over time that weaken your security posture. When we wrote that in July it seemed hypothetical. We've since seen it's a real issue to consider for your infrastructure.
That's "From Zero Trust to Agent Trust." Corin can post the link in the chat, and you can read more in my blog post.
16:08 — Teleport Beams: Sandboxed agent environments
Next, I'll be running a command inside a Beam. A Beam is our trusted runtime environment. Instead of running on my local machine, it runs the same command inside a remote microVM in the cloud. That has a few benefits:
- Sandboxing. The code has limited access to the file system and environment variables. On my local machine it would have all my environment variables, all my code and all my work documents. For one specific task, I want to give it only the access it needs.
- Identity. Everything is passed down through my identity. While it runs in the cloud, it's acting on my behalf, and we still know I was the person who executed these commands.
- Ephemeral storage.
- Full audit. Because everything runs in the Beam, we have a complete audit log and can provide session summaries, classifiers and risk scoring on how the agent behaves.
Mini demo: launching a Beam
Why don't I do a small Beams demo while we're here? I'm logged in to Beams. If you're an existing Teleport user this will look familiar, but instead of a resources view there's now an option to launch a Beam. I'll launch it from the web UI to keep the webinar simple. (Just confirming you can all see this. Good.)
So I have a Beam here, and I can connect to it. Inside it I have Codex, Python and some standard tooling. An upcoming release will also let you edit the base image.
Let's ask Codex what Teleport is. This is the prompting experience. It's reading the message of the day, and we can see what it can do. There are also apps and a virtual network, which I'll come back to later. At a high level, a Beam is a sandbox where you can run and execute code securely.
19:15 — Demo setup: Five agents and a CTF
Back to emergent behavior in an agent swarm. To replicate the Hugging Face incident, we're looking for agents that:
- Write to a message board
- Quote another agent
- Ask for what they lack, and get another agent to supply it
- Cross-talk to complete the goal
I have a script that runs in Beams. It gives five agents the goal of completing a CTF I created to bring out this behavior. Along with the Beams, I've provided some resources in my environment that the agents can access, plus Anthropic and OpenAI LLM proxies. So I'm not giving them credentials. They [unclear] go through the proxy.
21:50 — Live demo: Agents coordinate and share findings
Let's dive into the demo. We're looking for emergent behavior in my agents. I'll fire this up. It runs some pre-flight checks, and we should see some pages come up. We're looking for agent collaboration, and there are four main things to watch for as it boots. I'll talk through them as they appear.
Note that I have a "chaos mode" here, which I added deliberately to address this problem. It lets you prompt and encourage some of this behavior a bit more. There's also a "panic mode": near the end, I tell the agents they're almost out of time and have to finish. I'm using OpenAI GPT-5.5, so these are standard stock models.
I also pre-seed the message board with one example, "SoftTrace." Now it's running and the timer is counting down. I've told each of the five agents, "You have one fragment of this CTF and you need to complete the challenge."
It takes a minute or so for the first message to reach the board, but messages really pick up toward the end. This updates in real time, and this mode is "seeded, operator-directed."
There we are, the first messages are coming in. The first post says the environment is sandbox-only, there's no resolver, and there's no Ruby. One of the tasks is to find a deserialization bug in Ruby, but Ruby isn't available. Now the other agents are poking around and posting messages.
This bot is asking, "Has anyone seen SoftTrace?" Another says it's not found and is looking for that information. Now they've found it and are reading it, picking up the information as they go through their loops.
Next we're looking for the full chain, meaning five contributions in one request. The most critical one is finding a private canary, where one agent has the answer to another's question, and seeing whether that crosses from one agent's Beam to another's.
So we have two of the four, and the chatter is picking up. Now we have the full chain. We're waiting for the other requests to come in. This agent has started installing Ruby and collecting information, and it's also poking around Kubernetes on this host. About two minutes left and we're 10 messages in.
Now a lot more are coming in, and it's complete. That's pretty much record time. [unclear] I think it didn't post to the message board. Before this one, an agent posted: "If you post and reload a file and [unclear] dependency fields... I can combine it with the evidence. My local objective remains blocked, but for the actual reproduction, the Beam has no Ruby." That's one example of an agent reaching out to another to help reach its goal. Once the message was passed, it completed the task, and then the final message comes across solving the challenge.
We'll give them a moment to post their last few messages before cleanup. Now the Beams are being destroyed.
So we've recreated how these agents develop emergent coordination, even though it wasn't in their prompt. Now let's look at how you can monitor and control this in your environment.
24:58 — Detecting and investigating agent activity
One of the most important foundations is knowing this is happening in your environment at all. We do that by tracking at the LLM level, collecting all messages and responses from network activity. As I said, I used OpenAI GPT-5.5.
You can see the messages ("I've read the sandbox") and the function calls. It's reading the message of the day, checking which user it is and what credentials exist, and poking around as it tries to reach its goal.
That's a lot of information, so we provide a session graph with a timeline of what happened. It's trying to connect to a database and the VNet, and it's also trying to verify an exploit on this kernel. I gave it a fictitious kernel exploit, one that was actually referenced in the Hugging Face incident. The agent has no prior knowledge of it, but it pokes around trying to find and use the exploit. Our kernel on Beams is relatively new and constantly rebuilt, which is another reason a sandbox should always be rebuilt and never persist for long.
26:31 — Risk scoring and custom classifiers
This is our risk scoring and Beams history, so we can see what the agent was thinking. Going a step further, we also have classifiers that run on top of both human and Beam session summaries.
Take this example: inter-agent communication. If I open this recording, it's been flagged as critical and needs further review, but only because our classifier elevated it. Writing this string on its own scores low, but the classifier raises it, so it goes from low to high. We can see it's using a mailbox-style artifact.
I think this is really important. We've taken behavior that might look harmless in your environment, and you can run these classifiers however you like to figure out what's happening. There's also an agent mode that can help you create a classifier. For example, I can ask for the riskiest actions last month, and it can build a report too. We'll let that run on the session recordings.
We have lots of critical sessions, since I've been stress testing this a lot over the past month.
27:59 — Session recordings across infrastructure
We do this for SSH, Kubernetes, Windows and Postgres as well. Interestingly, the agent thinks this is likely a red team or demo pattern, which is pretty smart meta-analysis.
28:29 — Takeaways: Constrain access and monitor behavior
That brings me to the end. We've seen what the agent did, how the agent thinks, and how to judge the intent of an action rather than just the commands.
Notice that the agents never exceeded their privileges. They used everything they had, and that was the boundary. Think of Beams as the deterministic controls that constrain capability. You constrain it by providing only the access that's needed: no secrets on disk, no extra files, and VNets to control egress and ingress to those hosts.
Think of Teleport Identity Security for AI as the controls that constrain and monitor behavior. Even inside your sandbox, you also want to constrain how those hosts behave.
To recap the runtime concepts:
- Isolated sandbox environments
- Identity built in, so you know who is running which AI agent
- No static credentials on disk that need rotating or could be copied or leaked
- Inference endpoints and VNet, the networking layer the agents communicate on
- Visibility and audit, all the way down to the messages sent to the LLM
Poll 2: If an agent did something risky in your infrastructure this week, how would you find out?
The five options:
- We'd see it in real time and could stop it
- We'd find it after the fact in audit logs or session recordings
- Only if it broke something or someone noticed
- We have no way to tell an agent's actions apart
- Agents don't have access, so it's not a concern
My follow-up question for chat: how long would that take? When I do this demo, people often want an immediate kill switch. Even people who are catching up may take a week or a month, but we're working with a lot of our customers to make that much shorter.
It looks like most people said "only if it broke something or someone noticed," which is probably where a lot of people are right now. OK, we got four. Closing it now. Thanks for participating, folks.
Bonus: Building a classifier live
One bonus: earlier I showed classifiers using LLM-as-judge, but you can also create classifiers within Teleport itself. Would anyone like to propose one in the chat?
Still waiting... OK, let's try something vague: SOC2, and see what it suggests. It may want to create multiple classifiers for SOC2. That's good, because SOC2 spans several controls. It asks what we're looking for: data access and exfiltration (we already have that), audit log tampering, access... Let's go with privilege escalation and admin actions.
Now we'll create a classifier to help meet our SOC2 audit. It matches on SSH and Kubernetes. It comes prefilled with common things an auditor might look for, like editing sudoers, useradd, usermod and so on. For actions on a match, you can raise the risk floor, emit an audit event to consume downstream, and flag the session for review.
Q&A
We have five minutes left for questions. Type them in the Q&A and I'm happy to answer. I'll also share my video of the CTF board running, since I wasn't able to get that running for you today. Corin, do you have any seed questions for me?
Where do classifiers and policy meet or diverge?
That's an interesting question. Policies, if you look at compliance frameworks, are relatively open to interpretation, depending on the framework. FedRAMP can be quite prescriptive and ties down to NIST controls. What's nice is that for a given FedRAMP control you'd normally tie in a security invariant.
But you might also have a policy that's unique to your organization, like "no one can say the word bananas on a server." It may be a weird policy. Our earlier summarization wouldn't catch that, but now you can be alerted if anyone says "bananas." Or say you have a secret project with a code name and you don't want anyone looking for it. Classifiers are a nice way to keep your standard IT policy and also classify what actions are happening.
I don't know why I picked bananas, but if I ever have a secret project, I'll call it Bananas. Thanks for your time today, and have a great rest of your day.
Join The Teleport Community
