Rendered at 11:08:36 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
smashed 13 hours ago [-]
Coincidentally I mis-prompted claude code the other day while working on a toy project and failed to specify the project should be built on top of docker and not "like docker".
It went on to waste all my tokens creating a specialized docker clone. Cool I guess.
zoobab 1 hours ago [-]
I discovered proot-distro build yesterday, which does not require a special kernel with network and pid namespaces, it could run on older machines that do not have those features, or as a unix user that don't have those permissions.
The biggest change would be cgroup v2: one unified hierarchy where you just mkdir a group, set memory.max/pids.max and write the pid to cgroup.procs, instead of juggling per-controller mounts. clone3() with CLONE_INTO_CGROUP also removes the race of moving the child into the cgroup after it has already started running. Unprivileged user namespaces are enabled on most distros now, so a lot of it can be done rootless, and the new mount API (open_tree/move_mount) makes the rootfs setup less fiddly than the old pivot_root dance.
ranger_danger 17 hours ago [-]
> I wanted specifically to find a minimal set of restrictions to run untrusted code.
I don't think we should consider containers to be a security boundary. Even full VMs can be escaped, and have been, many times.
The fact that this is possible in the first place makes me think we need a much better approach.
chubot 17 hours ago [-]
As far as I know, Firecracker, gVisor, and Kata Containers are the solution here. They use VM primitives (x64_64 and ARM64 extensions) and have lighter codebases
But I don't have any direct experience with any of them. I'd be curious what people who have built on top of them think
edit: OK it looks like Kata can use Firecracker, so as far as isolation, it's either Firecracker or gVisor. And Firecracker is the VMM I mentioned, but gVisor is quite different -- it's more like a user space kernel that emulates syscalls.
binsquare 17 hours ago [-]
I'm going to toss in smolvm as well because firecracker needs some expertise to make the box usable and secure.
I've deploy gvisor, done basic test of firecracker and an honest attempt at production kata.
Firecracker and gvisor are nice systems not horrible to use, gvisor isn't quite the same security level though.
Kata is HARD to make. The technical know how to make that in production is awe inspiring. I wanted to use it but it was so complicated to integrate into a cluster I literally just gave up and mirrored raw VMs into the cluster which was alot easier actually.
Kata also breaks any potential of confidential VM unless you're a virtualization wizard.
You should go check out redhat's confidential container method for a production design overview. Their ARO self hosted system.
palata 12 hours ago [-]
> gvisor isn't quite the same security level though
Which one is more secure? I thought gvisor but your sentence sounds like it implies the opposite.
johnsmith1840 10 hours ago [-]
For my assumed goal of yours I would say none of these are good options. A simple VM or ec2 node hosting a coding sandbox is likely the better option.
All 3 of these are vastly different tech. Confusingly, you could run all 3 of these at the same time. I know that probally doesn't help. If you want to learn more I'd go get a linux machine somewhere like aws and play around with them.
A local agent harness is a "light" sandbox in of itself you'd also be able to learn alot from.
Bubblewrap is a good tech to check out and likely a better tool if you're looking for a simpler option. App armor in linux is also nice gvisor like tool.
palata 4 hours ago [-]
> For my assumed goal of yours
Wrong assumption, my goal was to get more information from someone who sounded knowledgeable and said "it's not quite the same security level", which was ambiguous to me.
> I know that probally doesn't help.
No, it doesn't. I know how to learn myself, I don't need someone to tell me that if I want to know more, I should go read about it.
imtringued 2 hours ago [-]
gVisor is basically a reimplementation of the kernel interface in go using a restricted subset of known to be secure system calls against the Linux kernel.
The underlying assumption behind it is that it is easier to build a secure kernel inside of golang than in C.
laurencerowe 17 hours ago [-]
As I understand it Kata supports multiple VMM backends, Firecracker, QEmu, Cloud Hypervisor, and their own Dragonball. Except QEmu, I believe those are all built on crates in the rust-vmm ecosystem, each making slightly different tradeoffs.
bityard 15 hours ago [-]
Depends on the threat model. Security is not black-and-white.
Containers protect against "I don't trust this curlpipe to not crap all over my dotfiles," rather than, "there might be a sandbox escape attack in this random file I downloaded."
If a VM is not sufficient for your threat model, I'm curious what is?
zamadatix 14 hours ago [-]
We don't have any "security boundaries" by this definition, just "security make-it-harder"s. I.e. "security boundaries" always have a relative strength associated with them, not a guarantee they keep the thing secure without any doubts.
SOLAR_FIELDS 12 hours ago [-]
The principle of defense in depth is built around the idea that with enough time, any system can likely be compromised, but the chance of compromising the system before being detected in your attempts to do so is lower the more safeguards you put into place
raesene9 16 hours ago [-]
I definitely wouldn't trust standard Linux style containers that expose a shared Linux kernel at the moment, there's been far too many LPE and container breakout vulnerabilities this year. It's possible that in future if the kernel gets a lot more hardened, that could change but things like Firecracker are a better bet from a security standpoint.
stryan 16 hours ago [-]
Podman supports using KVM backed virtualization for containers via libkrun: `podman run --runtime=krun` . Still not the end-all-be-all security boundary, but better I think.
LtWorf 14 hours ago [-]
They are a security boundary, but like everything else, not perfect.
mdspan 15 hours ago [-]
I think until something hardware-based like CHERI becomes widely deployed (which seems extremely unlikely in the near to mid term given), we're going to keep seeing VM escape CVEs pop up indefinitely.
LtWorf 13 hours ago [-]
Because we've never encountered hardware bugs…
mdspan 12 hours ago [-]
Critical hardware bugs occur an order of magnitude less frequently than critical hypervisor/kernel bugs, which is why they always make the news. In general, they're also more difficult to exploit. We haven't seen any serious spectre or meltdown malware in the wild almost a decade later.
LtWorf 4 hours ago [-]
But when they get discovered the only solution is to redesign the thing and throw away all of the hardware, with no mitigation strategy in between.
Some might prefer issues that can be fixed instantly and cheap to issues that will require years and millions to be fixed.
imtringued 2 hours ago [-]
You seem to be arguing on the basis that people are throwing away the software security layers, when the discussion so far has been about hardening the software security layers even more with secure hardware.
This is the nirvana fallacy in action. Because you cannot imagine a world with perfect hardware, you think hardware that catches 99.9% of security bugs is worthless.
aussieguy1234 5 hours ago [-]
I use per project bubblewrap scripts for my coding agents.
PunchyHamster 2 hours ago [-]
Container escape is already vanishingly small amount of attacks.
Like, if app in wild gets hacked, you kinda already are screwed, even if the hack is contained to the box (whether VM or container) app runs in, you still get whatever app keys app used, and you still get whatever visibility to internal network the app had. https://xkcd.com/1200/ basically.
If your app server gets hacked, all user data leaks anyway. If you divide everything to microservices so they see minimum required amount of data, attacker can still see everything that goes thru it
kragen 12 hours ago [-]
Last night I was looking for how to run Graphviz on untrusted input in a secure way, because recent versions of Graphviz give untrusted input to a whole insane rat's nest of code: Harfbuzz, Pango, libfribidi, libthai, libgraphite2, and its own internal format parser, each of which has a rap sheet of CVEs that makes Charlie Manson look like a petty shoplifter. And apparently Pango is even multithreaded, so we can expect nondeterminism. (Most of this doesn't show up in a simple ldd check; Graphviz sneakily waits until runtime to dlopen graphviz/libgvplugin_pango.so.6!) So, naturally, I wanted to sandbox it so that the worst thing a malicious attacker could do would be to make it draw Dickbutt or something. What I ended up with was less than 500 lines of code using Claude's suggestion of Bubblewrap http://canonical.org/~kragen/sw/dev3/wrapdot:
#!/bin/sh
# Confine dot in bubblewrap, taking input from stdin and writing PNG
# output to stdout.
# 64 megs seems to be enough, 21 megs isn’t.
address_space=64001000
# With zero --fsize, we can’t write the output file on stdout if it's
# redirected to a file, but you can pipe it to `cat`.
file_size=0
cpu_seconds=5
# Apparently Pango or fontconfig is multithreaded now‽
# (process:2): GLib-ERROR **: 00:37:18.600: creating thread '[pango] FcInit': Error creating thread: Resource temporarily unavailable
processes=4
# We’re using --unshare-user, etc., explicitly, because --unshare-all
# uses the wimpy --unshare-user-try and --unshare-cgroup-try options.
# --remount-ro / prevents malicious code from filling the root
# filesystem with empty files.
exec bwrap \
--ro-bind /bin /bin \
--ro-bind /lib /lib \
--ro-bind /lib64 /lib64 \
--ro-bind /sbin /sbin \
--ro-bind /usr/lib /usr/lib \
--ro-bind /usr/share/fonts /usr/share/fonts \
--ro-bind /var/cache/fontconfig /var/cache/fontconfig \
--ro-bind /etc/fonts /etc/fonts \
--remount-ro / \
--unshare-user --unshare-ipc --unshare-pid --unshare-net --unshare-uts \
--unshare-cgroup --die-with-parent --new-session --cap-drop ALL \
--clearenv --setenv PATH /bin \
prlimit --as="$address_space" --fsize="$file_size" \
--cpu="$cpu_seconds" --nproc="$processes" \
dot -Tpng -Gdpi=192
# For testing, to verify that network access is indeed blocked:
# nc.traditional -v -v 127.0.0.1 8000
Still, this is enough code that I'm not sure I haven't left something out. Still pending: run ImageMagick or netpbm inside the sandbox to convert the PNG file into a PPM or BMP — there have been CVEs in libpng in the past, and of course it's potentially vulnerable to zip bombs.
Of course this is still exposing most of the Linux kernel system call interface, although fortunately not /proc and /dev. And it's still fucking insane that drawing a node-link graph with three nodes requires more virtual memory than my first Linux machine had in total, in which it ran web browsers and recompiled the kernel. But that's a little further down the line.
drybjed 2 hours ago [-]
In Plan 9 you would get an easy way of namespacing by just unmounting some directories.
All those game mashups made recently by LLMs... Give me a port of Firefox on Plan 9, I'll be impressed then.
PunchyHamster 2 hours ago [-]
bwrap needs additive mode; take everything (aside some basics like stdin/out/err acces) by default, then add permissions.
It went on to waste all my tokens creating a specialized docker clone. Cool I guess.
https://pypi.org/project/proot-distro/
https://news.ycombinator.com/item?id=30623372 (250 points | March 10, 2022 | 27 comments)
https://news.ycombinator.com/item?id=22232705 (267 points | Feb 4, 2020 | 29 comments)
https://news.ycombinator.com/item?id=15608435 (440 points | Nov 2, 2017 | 53 comments)
https://github.com/nolabs-ai/nono
I don't think we should consider containers to be a security boundary. Even full VMs can be escaped, and have been, many times.
The fact that this is possible in the first place makes me think we need a much better approach.
https://firecracker-microvm.github.io/
https://gvisor.dev/
https://katacontainers.io/
But I don't have any direct experience with any of them. I'd be curious what people who have built on top of them think
edit: OK it looks like Kata can use Firecracker, so as far as isolation, it's either Firecracker or gVisor. And Firecracker is the VMM I mentioned, but gVisor is quite different -- it's more like a user space kernel that emulates syscalls.
https://github.com/smol-machines/smolvm
Firecracker and gvisor are nice systems not horrible to use, gvisor isn't quite the same security level though.
Kata is HARD to make. The technical know how to make that in production is awe inspiring. I wanted to use it but it was so complicated to integrate into a cluster I literally just gave up and mirrored raw VMs into the cluster which was alot easier actually.
Kata also breaks any potential of confidential VM unless you're a virtualization wizard.
You should go check out redhat's confidential container method for a production design overview. Their ARO self hosted system.
Which one is more secure? I thought gvisor but your sentence sounds like it implies the opposite.
All 3 of these are vastly different tech. Confusingly, you could run all 3 of these at the same time. I know that probally doesn't help. If you want to learn more I'd go get a linux machine somewhere like aws and play around with them.
A local agent harness is a "light" sandbox in of itself you'd also be able to learn alot from.
Bubblewrap is a good tech to check out and likely a better tool if you're looking for a simpler option. App armor in linux is also nice gvisor like tool.
Wrong assumption, my goal was to get more information from someone who sounded knowledgeable and said "it's not quite the same security level", which was ambiguous to me.
> I know that probally doesn't help.
No, it doesn't. I know how to learn myself, I don't need someone to tell me that if I want to know more, I should go read about it.
The underlying assumption behind it is that it is easier to build a secure kernel inside of golang than in C.
Containers protect against "I don't trust this curlpipe to not crap all over my dotfiles," rather than, "there might be a sandbox escape attack in this random file I downloaded."
If a VM is not sufficient for your threat model, I'm curious what is?
Some might prefer issues that can be fixed instantly and cheap to issues that will require years and millions to be fixed.
This is the nirvana fallacy in action. Because you cannot imagine a world with perfect hardware, you think hardware that catches 99.9% of security bugs is worthless.
Like, if app in wild gets hacked, you kinda already are screwed, even if the hack is contained to the box (whether VM or container) app runs in, you still get whatever app keys app used, and you still get whatever visibility to internal network the app had. https://xkcd.com/1200/ basically.
If your app server gets hacked, all user data leaks anyway. If you divide everything to microservices so they see minimum required amount of data, attacker can still see everything that goes thru it
Of course this is still exposing most of the Linux kernel system call interface, although fortunately not /proc and /dev. And it's still fucking insane that drawing a node-link graph with three nodes requires more virtual memory than my first Linux machine had in total, in which it ran web browsers and recompiled the kernel. But that's a little further down the line.
All those game mashups made recently by LLMs... Give me a port of Firefox on Plan 9, I'll be impressed then.