This is the mail archive of the libc-alpha@sourceware.org mailing list for the glibc project.


Index Nav: [Date Index] [Subject Index] [Author Index] [Thread Index]
Message Nav: [Date Prev] [Date Next] [Thread Prev] [Thread Next]
Other format: [Raw text]

Re: [RFC] Toward Shareable POSIX Signals


> On Fri, Mar 09, 2018 at 12:54:33PM -0800, Daniel Colascione wrote:
>> On 03/09/2018 12:25 PM, Rich Felker wrote:
>> >On Fri, Mar 09, 2018 at 02:30:51PM -0500, Zack Weinberg wrote:
>> >>>"Just use glib" is of course fundamentally unacceptable. But the
>> >>>obvious solution is "just use threads" and I don't see why that's not
>> >>>acceptable. The cost of a thread is miniscule compared to the cost of
>> >>>a child process, and threads performing synchronous waitpid can
>> >>>convert the result into whatever type of notification (poll wakeup,
>> >>>cond var, synchronous handling, etc.) you like.
>> >>
>> >>The main problem I see with this idea is, a thread waiting for _any_
>> >>process can steal the event from a thread waiting for a specific
>> >>process; this makes it nonviable for any situation where you don't
>> >
>> >I never proposed using a thread that calls wait or waidpid with a
>> >negative argument, rather one thread per child.
>>
>> Understood.
>>
>> >As long as there is no
>> >rogue thread in the program doing wait-any, the thread-per-child
>> >approach lets you emulate pdfork pretty well; programs written around
>> >this model can use pdfork as a drop-in replacement and eliminate the
>> >cost of the thread.
>>
>> My contention is that a thread per child process is infeasible from
>> a resource POV and that major subsystem authors will never adopt
>> this approach.
>
> This may be the current reality but my contention is that it's based
> on myths. A thread that will do nothing but waitpid can be created
> with a 1-page stack, no guard page, and all signals blocked. It
> consumes 4k of memory, 4k plus some epsilonish amount of kernel
> memory, one number from the pid/tid space (enlarge if needed), and a
> few microseconds to start/exit. Compare with exec'ing a child which
> takes hundreds of microseconds (fork-only is much less than with exec,
> but still much more than a thread, and fork-only should be considered
> deprecated for most purposes for lots of other good reasons). We're
> really talking about something like a 1-5% increase in cost here,
> probably on the lower end.

Myth or not, the idea that threads are expensive has a powerful hold on
people. Especially since thread creation _and teardown_ still contend on
mmap_sem in Linux, and unnecessary vm-map modifications are still highly
undesired.

Besides, everyone using threads for process waiting does nothing to help
*existing* software that uses broad wait operations, especially in the
case where we have a library that wants an internal helper child process
and that wants to work in arbitrary processes.

It's the logic that justifies O_CLOEXEC.

Besides, the kind of thread-based process waiting you're talking about is
pretty complex to implement, and wait is relatively simple and more
efficient. It's wait that people will default into using.

>> >>>On Fri, Mar 09, 2018 at 05:58:51PM +0100, Florian Weimer wrote:
>> >>>>But [threads] only works for asynchronous signals.  It's reasonable
>> >>>>for an application to want to catch synchronous signals (SIGBUS
>> >>>>when dealing with file mappings, SIGFPE for arithmetic), and there
>> >>>>is currently no thread-safe or library-safe way at all to do that.
>> >>>
>> >>>Yes, as I noted each use case needs to be considered separately to
>> >>>determine if there's some other better/more-portable/whatnot way it
>> >>>could be done already. The above applies only to SIGCHLD.
>> >>>
>> >>>FWIW I'm rather skeptical of many of the usage cases for synchronous
>> >>>signals (most are dangerous papering-over of UB for dubious
>> >>>performance reasons; never-taken "test reg,reg;jz" takes essentially
>> 0
>> >>>cycles on a modern uarch) but SIGBUS makes it hard to use mmap safely
>> >>>to begin with. So there's still a lot of material to consider here.
>> >>
>> >>If I remember correctly, GCJ tried to use signal handlers to generate
>> >>NullPointerExceptions not for speed reasons, but for code-size and
>> >>exception-precision reasons.  But it was never 100% reliable and it
>> >>might have been better to go with "test reg,reg;jz" + lean harder on
>> >>proving pointers couldn't be null.
>> >
>> >This is my view. Null checks/proofs should be maximally hoisted and
>> >explicitly emitted in the output rather than relying on traps.
>>
>> Every major managed code runtime team disagrees with you.
>>
>> It's not productive for low-level infrastructure maintainers to
>> claim that a universal practice is somehow illegitimate. This
>> attitude is not going to convince people doing the supposedly
>> illegitimate thing to stop doing it, but it will block progress that
>> leads to improvement of the system as a whole.
>
> There are a lot of widespread programming practices that have little
> of no legitimacy, and it is productive for parties who have some
> leverage to change them to try to use that leverage. Ideally this
> should not be unilateral (based on a single person's or single
> implementor's position) but reflect widely agreed upon principles.

First, relying on traps for optimization isn't an illegitimate technique.
There is _nothing_ wrong with it from a conceptual perspective. Legitimacy
comes from broad adoption. Traps in runtimes do work. They've worked well
and they've worked for a long time. They do improve performance. Why
should anyone stop using them?

Second, there's using leverage and there's tilting as windmills. The
objection to trapping is largely aesthetic; when performance bumps up
against aesthetics, performance has to win. There's no chance that major
managed-code runtimes leave performance on the table because some people
think trapping is ugly.

>> >>That's the only case I'm personally familiar with where a serious
>> >>application tried to _recover from_ synchronous signals.  I've also
>> >>dug into Breakpad a little, but that is a debugger at heart, and it
>> >>would be well-served by a mechanism where the kernel would
>> >>automatically spawn a ptrace supervisor instead of delivering a fatal
>> >>signal.  (This would also allow us to kick core dump generation out of
>> >>the kernel.)
>> >
>> >This is a very bad idea. Introspective crash logging/reporting is a
>> >huge gift to attackers. If an attacker has compromised a process in a
>> >manner to cause it to segfault, they almost surely have enough control
>> >over the process state to force the handler to perform code execution
>> >for them. There have been real-world CVEs along these lines.
>>
>> I've hacked on crash reporters for a while now. Reporting a crash in
>> a damaged process environment is undesirable, but unavoidable in
>> some cases. For example, on iOS, fork(2) doesn't work. At all.
>> Consequently, breakpad there needs to do its best with the state it
>> has.
>>
>> Calling fork(2) in a SIGSEGV handler and immediately execing a crash
>> reporting process is generally safe enough. It's hard for things to
>> go wrong enough that this mechanism doesn't work. That fresh crash
>> reporting process can ptrace its parent and collect what it wants.
>
> On i386, the vdso syscall pointer is stored at the beginning of the
> TCB, which is just above the thread's TLS, which is just above the
> thread's stack with no guard pages in between. Your syscall to fork
> could very well turn into a jump to the attacker's payload, not to
> mention all the other stuff done in addition to the fork.

Right. That's why you issue the system call _directly_ using something
like https://chromium.googlesource.com/linux-syscall-support/

>> While some kernel help in spawning this process wouldn't hurt, I
>> don't think it's particularly necessary. (And I think the existing
>> Linux core_pipe approach is adequate.)
>
> It's the only way to make it remotely secure. You cannot safely do
> anything from a compromised context. Any further logging/reporting
> work has to take place in a known-uncompromised context and has to
> account for any data structures extracted from the crashing process
> possibly being tainted/malicious.

Right. After a process has crashed, its entire address space is untrusted
input.

>
>> We _do_ need user-space dump collection though. The logic for
>> deciding what information we include in a crash report is too
>> complex to hoist to the kernel, where it'll seldom get updates. The
>> kernel's job should be limited to hooking up a crashing process and
>> a crash-reporting process; I'd get rid of kernel-written core dumps
>> entirely if I had my way.
>
> This is plausible if you sandbox the collection utility such that it
> does not have access to do harmful things locally and does not have
> channels for exfiltration.

Agreed. I've implemented such sandboxed collection systems.


Index Nav: [Date Index] [Subject Index] [Author Index] [Thread Index]
Message Nav: [Date Prev] [Date Next] [Thread Prev] [Thread Next]