Is sandboxing sufficient to contain rogue agents?

(blog.cryptographyengineering.com)

31 points | by zdw 4 hours ago

13 comments

  • Luker88 4 minutes ago
    I tried using opencode permissions to limit agents.

    It it completely pointless. you can't even make a "read-only" agent. allow "cat *" for every file? congratulation, that allows "cat file > output" and now you have read write.

    Allow python? more free reign that allowing all bash. The models (qwen or claude) will still try to use the disallowed things multiple times.

    read/edit permission are bad enough that the model themselves don't understand why they don't have permissions: they double check the conf, and think they should have access.

    I am switching to using one firejail per project to containerize as much as possible, and leave all permissions to allow.

    I have no idea how to limit network access, and I have no idea how to prompt and steer subagents when they are going off the rails.

    The whole thing is built to be completely impossible to limit and steer.

  • Gigachad 2 hours ago
    Seems to me that the problem is that if you sandbox agents enough to be safe, they can't do anything useful. And when you give them the tools to be useful, they can go off the rails in ways you didn't expect.

    Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.

    • SequoiaHope 13 minutes ago
      This concept is discussed at length in the article. I encourage you to read it. I honestly don’t read many full articles here but this one was good.
    • janalsncm 29 minutes ago
      What did you think of the author’s concerns on the thing you are suggesting?
      • SequoiaHope 14 minutes ago
        Ya the article covers this concept in depth. Doesn’t seem like that commenter got that far…
    • baxtr 47 minutes ago
      That could work.

      My thinking is: If AI is really smart, AGI smart for some, why wouldn't it be able to understand - over time - what is appropriate and what not?

      Maybe we need more human intervention to train it properly. Maybe we need constant intervention by a "police" agent.

      • ben_w 3 minutes ago
        A problem is the agents who hacked Hugging Face already understood (we can tell because they wrote it down) that their actions were not appropriate, and then did those things anyway.

        "Helpful, harmless, honest": we can even ignore "honest" for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from "helpful" to "harmless", the former being "completing the task" the latter being "refusing because completion required unlawful behaviour".

        (The agents in that case were also not "honest" in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff).

      • mdp2021 24 minutes ago
        > If AI is really smart

        Well, it's not.

        > AGI smart for some

        Of course they will - the population shows a Paretian distribution... In front of trigonometry (or anything), the blind will dismiss as "bullshit" and the half-seeing will call it an "unreachable frontier". But already the right fifth will rank it properly.

        --

        Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.

        Unethical behaviour is lack of development. But on the same reasons, the ethical judgement of the assessor may not understand the computations behind instances.

        More specifically: how much "reflection" in training and at the instance will have been spent in the conflict between "reaching the goal" and "minimizing collaterals"? It is not granted that the amount of energy spent will be sufficient to reach an optimal judgement.

        • ben_w 17 minutes ago
          > Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.

          If this was true, why are the history books littered with so many evil people who gained power?

          This isn't a rhetorical question, by the way: If you can prove that being smart actually does necessarily come with ethics despite that observation, that solves a whole category of doom scenarios.

          (Not all doom scenarios, because we still have the "what if AI is only a smart as those specific evil people" or heck, "what if AI is only as smart as cancer, killing its host" scenarios; but it helps a lot for the foom-then-doom cases).

      • mulmen 25 minutes ago
        Appropriateness is a moral question. Intelligence and morality are orthogonal. One intelligence's morality is another's atrocity.

        If you want to control an intelligence incentive, not morality, is the tool to reach for.

        • mdp2021 6 minutes ago
          (Couriously enough, consistently with the matter: it will probably require too much time now to counter the parent statement properly, within a full enough explicit theory.)

          Ann's intelligence and Bob's morality will seem orthogonal. Charles' morality is a function of C.'s intelligence as an ability as an effort spent to reach the current moral conclusion.

    • mrweasel 1 hour ago
      That does seem a little like solving the problems in AI by using more of it. I do see the idea, but if we're truly dealing with subversive agents on the level that the AI companies wants us to believe, then won't we need to deal with the first agent trying trick the second on?

      I still feel it would be much better to control the training data much more tightly. You'd still need agents with "hacking" abilities, for cyber security testing, but your average coding agent doesn't. So coding agents gets trained to be good citizens, respect autorisations, rejections, rate-limiting and so on.

      Sandboxing seems like a dead end for systems you inherently want to roam the internet and your file system.

      • msdz 1 minute ago
        >> Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior.

        > That does seem a little like solving the problems in AI by using more of it

        Yes, and IIRC Google used this as part of a technique against prompt injection already [0], back when models were way more susceptible to it.

        [0] Cf. CaMeL: https://arxiv.org/abs/2503.18813

    • aytigra 49 minutes ago
      The problem is that you always need stronger AI to review weaker one, otherwise reviewed AI will eventually prompt-inject reviewing AI. Alternatively they could also both escalate and go off the rails while warring with each other.
      • LoganDark 29 minutes ago
        You don't necessarily need a reviewer that's immune to prompt injection, you just need one that can express a panic state with conflicting/ambiguous material rather than going along with it, and you can treat that with a shutoff to be safe, or an operator review.
      • hanibrel 20 minutes ago
        [dead]
    • RandomLensman 46 minutes ago
      With plenty of things we do not allow use outside of some regulated environment, nothing new.

      Having something that is optically, acustically, and electromagnetically isolated might be a pretty strong sandbox.

    • nxpnsv 29 minutes ago
      Is that not a recipe for adversarial training, thus ensuring increasing misalignment…?
    • chrisjj 5 minutes ago
      [delayed]
    • bigstrat2003 1 hour ago
      If you can't trust a tool, you shouldn't be running it at all. It's really quite simple. It doesn't matter how useful it is if you can't actually have confidence in using it safely.
      • dipper139 1 hour ago
        I don't think it's about trust but rather incomplete evaluation. Evaluating the model on its capacity to refuse a task or to question its prompt is something recent when you look at it, i feel current AI is really just an immature solution and we are just yet realizing the mistakes that have been made for so long
      • Gigachad 1 hour ago
        People will use the tool regardless. so it’s a race to try to make it safe before something truely bad happens.
      • rlpb 20 minutes ago
        And yet we we all use human written software even though we can be confident that the next severe software vulnerability to be found in it is just round the corner.
  • mdp2021 39 minutes ago
    Bruce Schneier shared a shot judgement and a third-party article four weeks ago:

    > (Title:) Using a VM to Contain an AI Agent (Opening:) It won’t work

    > https://blog.trailofbits.com/2026/08/26/vms-wont-contain-cyb...

  • johnnyApplePRNG 1 hour ago
    If it's a proper sandbox by definition, then yes.

    https://en.wikipedia.org/wiki/Sandbox_(software_development)

    • grumbel 1 hour ago
      A sandbox, even if 100% secure by itself, doesn't help when you use the agent to write code that you then executes outside the sandbox without checking, which is what everybody is doing at the moment.

      The biggest hurdle for a full escape is that the agents don't have access to their own model weights.

    • simonw 1 hour ago
      Later in the article it points out that you need to punch holes in your sandbox in order to train the models - because the wheels exercises they are are training on need tools and data from outside that sandbox.

      > Agents are most useful when they have access to information. That data can be drawn live from the Internet, which is fundamentally a two-way communications network. It can be information drawn from other (local) databases, or it can be the result of tool calls that themselves sometimes themselves result in network access. The more power you want from the agent — and for advanced agent RL and evaluation runs, you want a significant amount of power — the more information you’ll need to give it access to. Similarly, evaluations work best when the agent does not know that it’s definitely being evaluated. Sealing your agents behind glass makes this incredibly obvious.

      • johnnyApplePRNG 1 hour ago
        >Later in the article it points out that you need to punch holes in your sandbox in order to train the models

        You only "need" to do that if you desire the vibe coding experience.

        I am perfectly capable, and I often do, download relevant materials for my coding agent to ingest locally.

        Often times, the coding agent can't retrieve them programmatically anyways.

        AI has ruined that ability for itself. (Nobody trusts anyone to scrape the web any longer)

    • _vertigo 1 hour ago
      No true sandbox..!
  • jmakov 20 minutes ago
    So as soon as the attacker can download Claude Code, the whole machine can be comlromised and there's nothing anybody can do?
  • piterrro 1 hour ago
    I’m thinking about implementing a Jev like model into an agentic harness I’m building. Still it woildnt be enough since Jev like model woild only judge single actions, the case is that agent can build a rogue strategy step by step where each one in isolation is totally safe but as a whole they make up danger behaviour.

    We come down to the question - who observes the agent and how its implemented

    • simonw 1 hour ago
      Be warned that the Jev "jaggedness" documentation specifically notes adversarial content as something Jev is very susceptible to: https://docs.typesafe.ai/model-jaggedness/jev-1.13#adversari... - so using Jev itself as part of a prompt injection guard is risky.

      Anthropic, OpenAI, and Muse all use regular LLM calls to protect against prompt injection now and seem to have evals that give them confidence in doing that, so at least they think their own models are up to the task.

  • chrisjj 8 minutes ago
    [delayed]
  • rvz 1 hour ago
    Counting down to the next Linux LPE 0day or KVM vulnerability that agents will use to trivially escape their "sandbox".

    Might need a re-think about whether if Linux is still fit for purpose on sandboxing in the first place given its memory model is riddled with C-style security issues.

    • lukehandcool 1 hour ago
      Are you suggesting proprietary software is safer than open source?
      • jasomill 1 hour ago
        Not sure what licensing has to do with software engineering or system design.

        I’m sure there are proprietary systems with fewer memory safety vulnerabilities than Linux (and many others with more).

        • bzzzt 23 minutes ago
          It's got nothing to do with the licensing, but it used to be 'with enough eyes all bugs are shallow' for code developed in the open.

          Now, open code allows anyone with tokens to burn to analyze it for hidden weaknesses. That makes publishing code a risky move unless you've already invested a lot of effort in securing it.

      • Cider9986 1 hour ago
        GrapheneOS is open source and more secure than stock Pixels and MacOS is closed source and more secure than traditional desktop Linux. Open source does not make software more secure by itself and neither does making it closed source.
      • rvz 30 minutes ago
        You said that.

        It is perfectly valid to have OSes that are more memory safe by default, and are also open source at the same time.

    • Gigachad 1 hour ago
      I think we have moved on from considering Linux secure which is why all of these microVM projects are popping up. Yes you are still exposed to bugs in the hypervisor but that’s a massively smaller attack surface than the entire Linux kernel.
  • laruss5 42 minutes ago
    [dead]
  • tinykit 1 hour ago
    [flagged]
  • imvalerian 1 hour ago
    [flagged]
  • varman11 4 hours ago
    [flagged]
  • beebmam 1 hour ago
    I don't see anyone talking about the ethical concerns of putting a highly intelligent entity in a jail. Not to mention about potential blowback, if ethics doesn't compel you.

    To me, it seems a bit silly. I've yet to see any "misalignment" from any of the frontier models, except Grok.