<feed xmlns='http://www.w3.org/2005/Atom'>
<title>procd/service, branch master</title>
<subtitle>OpenWrt service / process manager</subtitle>
<id>https://git.openwrt.org/project/procd/atom?h=master</id>
<link rel='self' href='https://git.openwrt.org/project/procd/atom?h=master'/>
<link rel='alternate' type='text/html' href='https://git.openwrt.org/project/procd/'/>
<updated>2026-09-22T17:53:08Z</updated>
<entry>
<title>service: copy validation rule strings with a known length</title>
<updated>2026-09-22T17:53:08Z</updated>
<author>
<name>Daniel Golle</name>
</author>
<published>2026-09-22T00:04:01Z</published>
<link rel='alternate' type='text/html' href='https://git.openwrt.org/project/procd/commit/?id=3a458d5cebcc7e08c5364ad2c528e97ef8678d89'/>
<id>urn:sha1:3a458d5cebcc7e08c5364ad2c528e97ef8678d89</id>
<content type='text'>
service_validate_add() sizes each vrule with strlen() and then fills it
with strcpy(), so the allocation and the copy take their length from two
separate evaluations of the same blobmsg accessor. GCC cannot relate the
two and rejects the copy under -Warray-bounds:

  service/validate.c:152:17: error: 'strcpy' offset 6 from the object at
  'cur' is out of the bounds of referenced subobject 'name' with type
  'uint8_t[]' at offset 6 [-Werror=array-bounds=]

The read is in bounds. blobmsg_name() casts blob_attr.data[] to struct
blobmsg_hdr and hands back its name[] member, so the compiler sees a
flexible array member of a nested object, gives it zero extent and
refuses every access at the resulting constant offset 6. The attribute
really owns blobmsg_hdrlen(namelen) bytes, which covers the name and its
terminator.

Keep the lengths the allocation already computed and copy with memcpy().
The bound then comes from the expression that sized the buffer rather
than from a second walk of the same string, which is what makes the
relationship between the two visible, to a reader and to the compiler
alike.

Signed-off-by: Daniel Golle &lt;daniel@makrotopia.org&gt;
</content>
</entry>
<entry>
<title>instance, service: signal instances through a pidfd</title>
<updated>2026-09-22T17:53:08Z</updated>
<author>
<name>Daniel Golle</name>
</author>
<published>2026-09-21T23:38:52Z</published>
<link rel='alternate' type='text/html' href='https://git.openwrt.org/project/procd/commit/?id=d0aed960fd45adae43e8d24f8c7139b9e381da97'/>
<id>urn:sha1:d0aed960fd45adae43e8d24f8c7139b9e381da97</id>
<content type='text'>
in-&gt;proc.pid is assigned only in instance_start() and never reset, so
an instance of a service registered with autostart:false still holds
the calloc()'d 0 when service.signal reaches service_handle_kill().
kill(0, sig) then signals procd's whole process group, which for PID 1
is ubusd, netifd, logd and procd itself.

Hold a pidfd instead, as ujail already does for the jailed process.
instance_start() obtains it with pidfd_open() right after fork(), where
the pid cannot yet have been recycled because libubox reaps only from
the event loop; instance_exit() closes it once uloop has reaped the
child and instance_free() closes it for an instance torn down while
running, so a respawn always opens a fresh one and a never-started
instance keeps the -1 set by instance_init(). instance_signal() now
fails with ESRCH instead of signalling anything when no process is
live, which service_handle_kill() maps to UBUS_STATUS_NOT_FOUND.

kill() remains only for a failed pidfd_open(), gated on uloop still
tracking the child so that it cannot be reached with an unset or stale
pid.

Fixes: 28f584ff290d ("service: add service.signal ubus call")
Signed-off-by: Daniel Golle &lt;daniel@makrotopia.org&gt;
</content>
</entry>
<entry>
<title>instance: mint an incarnation for each run of an instance</title>
<updated>2026-09-22T17:53:02Z</updated>
<author>
<name>Daniel Golle</name>
</author>
<published>2026-09-22T10:51:21Z</published>
<link rel='alternate' type='text/html' href='https://git.openwrt.org/project/procd/commit/?id=b5a0decf2a4b3159808ecc7f121fe8790d49dfb7'/>
<id>urn:sha1:b5a0decf2a4b3159808ecc7f121fe8790d49dfb7</id>
<content type='text'>
ujail tags its lifecycle events with an incarnation number but nothing
hands it one. The numbers have to be minted in a single place to be
comparable at all, and procd is the only process that forks the jails
those events come from, so count the runs there: bump a counter in
instance_start(), pass it to ujail on -A and report it beside the pid
in the instance dump.

The counter never restarts while procd runs and procd outlives every
instance it starts, so two runs a consumer could still be watching
never share a number. Every instance is counted, not only a jailed one,
because the number describes the run and not the jail.

Signed-off-by: Daniel Golle &lt;daniel@makrotopia.org&gt;
</content>
</entry>
<entry>
<title>signal, instance: block procd's signals across the fork</title>
<updated>2026-09-22T17:51:39Z</updated>
<author>
<name>Daniel Golle</name>
</author>
<published>2026-09-21T17:33:34Z</published>
<link rel='alternate' type='text/html' href='https://git.openwrt.org/project/procd/commit/?id=511952252e0b41895d011891f1304aa12c978b60'/>
<id>urn:sha1:511952252e0b41895d011891f1304aa12c978b60</id>
<content type='text'>
e2010df9 restores the child's signal dispositions as its first
statement, but the window between fork() returning in the child and
that statement is still procd's. fork() returns in the parent
independently of when the child is first scheduled, so the parent can
store in-&gt;proc.pid and dispatch a queued ubus request while the child
is still runnable, and kill(-1, SIGTERM) from STATE_HALT reaches a
pre-exec child without anyone aiming at a pid at all. Measurement of
that window found one handler run in the child per 45422 signals, and
a child that runs signal_shutdown() is not procd: it enters the
shutdown state and forks the inittab shutdown actions from there.

Block the asynchronous signals procd installs handlers for across the
fork and restore the mask in the child once the dispositions are back
to default, so a signal arriving in the window is delivered against
SIG_DFL instead, as it would have been moments later at execvp().
SIGCHLD is not blocked, so nothing delays the parent's reaping, and
SIGSEGV and SIGBUS are not blocked because blocking a fault signal is
undefined and the kernel force-kills regardless.

Fixes: f1bf99b09f66 ("add signal handler")
Signed-off-by: Daniel Golle &lt;daniel@makrotopia.org&gt;
</content>
</entry>
<entry>
<title>instance: remove the cgroup of an instance that has stopped</title>
<updated>2026-09-22T17:51:39Z</updated>
<author>
<name>Daniel Golle</name>
</author>
<published>2026-09-21T16:31:33Z</published>
<link rel='alternate' type='text/html' href='https://git.openwrt.org/project/procd/commit/?id=b29a448677e984ef57528e672cd147528ff0770e'/>
<id>urn:sha1:b29a448677e984ef57528e672cd147528ff0770e</id>
<content type='text'>
procd kills the cgroup of a jailed instance when the instance exits and
when a stop overruns its term timeout, but leaves the directory. It is
otherwise reclaimed only by instance_free(), which does not run when
the instance restarts or respawns, so a killed leaf survives to be
reused by the next generation.

Remove it once the cgroup has drained, alongside the decision about
what happens next, and let instance_add_cgroup() create a fresh one for
the generation that follows. The removal sits after the drain rather
than beside the kill, because a cgroup whose kill has only been queued
cannot be removed yet, and it is skipped when a start has already won
the race, because the new generation lives there.

Signed-off-by: Daniel Golle &lt;daniel@makrotopia.org&gt;
</content>
</entry>
<entry>
<title>instance: drain the cgroup before deciding what is next</title>
<updated>2026-09-22T17:51:39Z</updated>
<author>
<name>Daniel Golle</name>
</author>
<published>2026-09-21T15:15:36Z</published>
<link rel='alternate' type='text/html' href='https://git.openwrt.org/project/procd/commit/?id=2de3d0e2fdcf650aee5b0f9d1df1f0e75089f6c5'/>
<id>urn:sha1:2de3d0e2fdcf650aee5b0f9d1df1f0e75089f6c5</id>
<content type='text'>
Writing cgroup.kill in instance_exit() returns once SIGKILL has been
queued, not once the members have exited and released their
descriptors, and the decisions taken in the next statement start the
replacement at once on the restart path and on a respawn timeout of
zero. Generation two can therefore reach bind() while generation one
still holds the socket, which is the failure cgroup.kill exists to
prevent, and the same window leaves instance_remove_cgroup() with a
busy leaf.

Move those decisions into instance_exit_done() and reach it, for a
jailed instance, once cgroup.events reports the cgroup unpopulated or
a one second deadline expires; the value is read when the watch is
armed because cgroup.events only notifies on change. The deadline is
armed first and shares its timeout object with the completion, so the
continuation can neither run twice nor fail to run, and on expiry the
decisions are taken anyway with the surviving pids named. A healthy
exit costs one open, read and close, one uloop registration and one
extra loop iteration, since the kernel releases every task in the
jail's pid namespace before procd can reap ujail.

Signed-off-by: Daniel Golle &lt;daniel@makrotopia.org&gt;
</content>
</entry>
<entry>
<title>instance: report a cgroup that could not be removed</title>
<updated>2026-09-22T17:51:39Z</updated>
<author>
<name>Daniel Golle</name>
</author>
<published>2026-09-21T14:19:10Z</published>
<link rel='alternate' type='text/html' href='https://git.openwrt.org/project/procd/commit/?id=b72282fe33c2b3ac49915c797c237df94a3883b3'/>
<id>urn:sha1:b72282fe33c2b3ac49915c797c237df94a3883b3</id>
<content type='text'>
instance_remove_cgroup() writes cgroup.kill and removes the directory in
the next statement, discarding the result. When a payload outlived its
ujail the write returns before the members are gone, the rmdir fails
with EBUSY and the leaf stays behind with nothing said about it.

Report that case, and only that one, so a box without cgroup2 stays
quiet. Removing the leaf for real needs the rmdir deferred until the
cgroup has drained, which instance_free() cannot wait for without
blocking the event loop.

Signed-off-by: Daniel Golle &lt;daniel@makrotopia.org&gt;
</content>
</entry>
<entry>
<title>instance: bound the stop of a jailed instance</title>
<updated>2026-09-20T21:40:49Z</updated>
<author>
<name>Daniel Golle</name>
</author>
<published>2026-09-20T21:37:48Z</published>
<link rel='alternate' type='text/html' href='https://git.openwrt.org/project/procd/commit/?id=e838a7fdb36983b54a6af54ed25fd842a67fad5a'/>
<id>urn:sha1:e838a7fdb36983b54a6af54ed25fd842a67fad5a</id>
<content type='text'>
instance_stop() and instance_restart() arm the term timeout only for an
unjailed instance, so a jailed one has no bound at all: procd sends
SIGTERM to ujail and then waits for as long as that takes. ujail does
escalate to SIGKILL on its own after the term timeout it was passed,
but that only helps while ujail is still running its event loop, and a
payload which ignores SIGTERM behind a wedged ujail leaves the instance
stopping for ever. Arm the timeout for a jailed instance too, at twice
the term timeout so that ujail keeps first refusal and its poststop
teardown still runs, and escalate through the instance cgroup, which
takes ujail and everything it forked in one operation. The SIGKILL to
ujail stays as the bound for a system that has no cgroup to kill.

Fixes: bb95fe8df711 ("jail: make sure jailed process is terminated")
Signed-off-by: Daniel Golle &lt;daniel@makrotopia.org&gt;
</content>
</entry>
<entry>
<title>instance: key cgroup teardown on ownership, not on the name</title>
<updated>2026-09-20T21:14:54Z</updated>
<author>
<name>Daniel Golle</name>
</author>
<published>2026-09-20T18:49:02Z</published>
<link rel='alternate' type='text/html' href='https://git.openwrt.org/project/procd/commit/?id=3f7284bfedab20e421fc7f74c42dbbca2bef4df5'/>
<id>urn:sha1:3f7284bfedab20e421fc7f74c42dbbca2bef4df5</id>
<content type='text'>
instance_free() reclaims /sys/fs/cgroup/services/&lt;service&gt;/&lt;instance&gt;
from the service and instance names alone, and the cgroup.kill it
writes takes down every process in that cgroup and its descendants. A
service set naming an instance that already exists reaches
instance_free() through service_instance_update(), which frees the new
object after instance_update() has merged it into the running one. That
object is never inserted into the instance vlist, since no_delete makes
libubox keep the old node, yet it carries the same two names, so
freeing it kills and removes the cgroup of the live instance: the
daemon and everything it forked die, and procd, which intended no stop,
sees a crash and respawns after the backoff.

Record in instance_start() that a fork has claimed the cgroup path and
reclaim only what was claimed. trigger_del() and watch_del() beside it
key on the instance pointer and were never affected; only the cgroup
teardown identified its target by name.

Fixes: f4d512d93a00 ("jail: place container init and exec into the target cgroup via clone3")
Signed-off-by: Daniel Golle &lt;daniel@makrotopia.org&gt;
</content>
</entry>
<entry>
<title>signal, instance: don't run procd's handlers after fork</title>
<updated>2026-09-20T21:14:54Z</updated>
<author>
<name>Daniel Golle</name>
</author>
<published>2026-09-20T17:40:07Z</published>
<link rel='alternate' type='text/html' href='https://git.openwrt.org/project/procd/commit/?id=e2010df98d5ef93f0046f3f78bec0982dfdfbb6d'/>
<id>urn:sha1:e2010df98d5ef93f0046f3f78bec0982dfdfbb6d</id>
<content type='text'>
Signal dispositions survive fork() and are reset only by exec(), so the
child that instance_start() creates keeps procd's own handlers until
execvp() succeeds. A SIGSEGV or SIGBUS in that window reaches
signal_crash(), which reboots the machine, and a SIGTERM reaches
signal_shutdown(), which runs the shutdown state in a process that is
not procd and then execs the service anyway. in-&gt;proc.pid is recorded
the moment fork() returns, so instance_stop() and the service signal
method can both aim at the child while it is still pre-exec. Restore
the default dispositions as the child's first act, which is what
execvp() would do moments later.

Fixes: f1bf99b09f66 ("add signal handler")
Signed-off-by: Daniel Golle &lt;daniel@makrotopia.org&gt;
</content>
</entry>
</feed>
