14 Commits

Author SHA1 Message Date
xfy
832b7c626d fix(infra): Docker daemon 不可用时不再 panic
DOCKER_CLIENT 是 LazyLock<Docker>,socket 缺失时 .expect() 直接 panic;
而 release profile 是 panic="abort",会让整个服务进程挂掉——但博客本身并不依赖 Docker。

改为 LazyLock<Option<Docker>>:连接失败记 error 日志后存 None。
新增 get_docker() 在 None 时返回 IOError(NotFound),run_in_container /
run_in_container_stream 通过 ? 向上冒泡,最终在 execute.rs 的错误脱敏层映射为
「系统暂时不可用」,其余功能不受影响。
2026-07-15 17:23:56 +08:00
xfy
65a7b1226f style: 全项目格式规范化(cargo fmt)+ Docker daemon 断连容错
- 全项目 .rs 统一 cargo fmt 排版(换行/缩进/参数分行/导入排序)
- infra/docker: DOCKER_CLIENT 连接失败不 panic,降级返回 None
  (博客不依赖 Docker,panic=abort 下会导致整个进程崩溃)
- infra/mod: 去除多余空行
- database/mod: 模块声明按字母序重新排序
2026-07-15 17:22:06 +08:00
xfy
e18d2bfb2c fix(docker): 用 tmpfs mode=1777 替换 uid/gid 选项以兼容 Podman 2026-07-14 16:42:04 +08:00
xfy
cc4c551aaa fix(code-runner): 容器清理失败重试+日志告警,避免静默泄漏
ContainerGuard::drop 之前用 let _ = 吞掉 remove_container 的错误,
Docker daemon 瞬时不可用会让容器泄漏且无人知晓。

改为最多重试 3 次(指数退避 200/400/800ms),仍失败则记录 error 级日志
(带 container_id 便于 docker rm -f 兜底)。容器清理是 fire-and-forget,
无法把错误回传业务层,所以用日志暴露故障而非 panic。
2026-07-13 17:56:42 +08:00
xfy
e50941ef24 fix(runner): propagate duration_ms through SSE done event
The SSE streaming path hardcoded '耗时: -' because duration_ms was
never threaded through:

- docker.rs: OutputChunk::Done now carries duration_ms, computed from
  start_container to wait completion via Instant::elapsed
- sse.rs: DonePayload gains duration_ms, serialized into the done event
- runner.rs: WASM done callback formats the real duration instead of '-'

The polling fallback path already had correct duration via ExecResult;
only the SSE path was missing it.
2026-07-13 10:09:15 +08:00
xfy
2059a55e11 fix(runner): make wait_container concurrent with log reading
run_in_container_stream was sequentially awaiting wait_container before
reading the log stream. wait_container blocks until the container
exits, so by the time the log loop ran, all output was already buffered
in the attach stream and got drained near-instantly — appearing as a
one-shot dump instead of streaming.

Now uses tokio::select! to run log_reader and wait_with_timeout
concurrently:

- log_reader: drains the attach stream, pushing each chunk via tx
  immediately as it arrives
- wait_with_timeout: waits for container exit (with timeout, kills on
  timeout)

Whichever finishes first wins; the survivor is drained/handled. When
wait completes, any tail bytes still in the stream are flushed before
sending the Done chunk.

This is the root cause of the streaming not working — a sleep(1) loop
now streams line-by-line instead of dumping all at once after exit.
2026-07-13 09:59:03 +08:00
xfy
832dd756b9 feat(runner): add streaming execution path in docker.rs
run_in_container_stream: same container lifecycle + ContainerGuard
cleanup as run_in_container, but pushes output chunks to an mpsc
Sender as the log stream is read, instead of buffering until the
container exits. Also retains a full buffer for the caller to write
back to EXEC_TASKS (polling fallback path).

- OutputChunk enum: Stdout/Stderr/Done{exit_code,oom_killed,timed_out}
- client disconnect detection: tx.send fails → stops pushing but
  continues draining the log stream so the container exits cleanly
- Done chunk carries terminal status for the SSE done event
- timeout/inspect/OOM logic identical to run_in_container
- run_in_container and its 4 tests left untouched
2026-07-10 11:41:37 +08:00
xfy
1584e01425 chore(deps): update all cargo and pnpm dependencies to latest versions 2026-07-09 18:12:54 +08:00
xfy
0ab3340133 fix(code-runner): allow exec on /tmp tmpfs for compiled languages
Docker tmpfs 挂载默认带 noexec,编译型语言(go/rust)把编译产物
落在 /tmp 后再 exec 会报 permission denied:
  - rust: /tmp/main (run-rust.sh 的 rustc 输出)
  - go:   /tmp/go-build*/b001/exe/main (GOTMPDIR=/tmp 的链接器输出)

解释型语言(python/node)执行根文件系统里的解释器,不在 tmpfs 上,不受影响。

给 /tmp tmpfs 显式加 exec 选项。这是容器执行的固有需求——只要
沙箱支持任何编译型语言,/tmp 就必须可执行。安全模型由只读根文件系统 +
cap_drop=ALL + no-new-privileges + pids/memory/cpu 限制兜底,/tmp 可执行
不引入额外攻击面(容器内本就允许执行自身进程)。
2026-07-09 11:04:27 +08:00
xfy
a7a97fc299 fix(code-runner): 去掉 nproc ulimit,修复容器启动 EAGAIN
容器启动时报 'exec /bin/sh: resource temporarily unavailable',根因是
HostConfig 里的 ulimit nproc=64 与 non-root user(1000:1000) 组合:
RLIMIT_NPROC 按 UID 计数且在容器初始 exec 阶段就生效,与容器内实际进程数无关,
导致 sh 都起不来。pids_limit=64 已在 cgroup 层兜底进程数限制,nproc 是冗余且
有害的双重约束。

复现:docker run --ulimit nproc=64 --user 1000:1000 → exec 失败
       去掉 nproc(保留 nofile + pids_limit + 其他沙箱约束)→ python/node 正常运行

cargo test infra::docker 全过。
2026-07-05 22:05:01 +08:00
xfy
f3eb93f320 docs(agents): document code runner and fix clippy/biome lint
文档:
- AGENTS.md 新增「Code Runner」架构章节(三层架构 + WASM 可见性规则 +
  governor 0.8 无 per_day 的处理 + runner 镜像契约)
- Environment 段补齐 CODE_RUNNER_* 与 RATE_LIMIT_CODE_EXEC_* 环境变量

Lint 修复(clippy --all-targets --all-features -D warnings 全绿):
- runner.rs/runner.rs: use_signal 冗余闭包 → 直接传 fn
- runner_config.rs: &*RUNNER_CONFIG 去 deref;assert_eq!(...false/true) → assert!
- markdown.rs: 局部变量 /// 改 //;Option.map().flatten() → and_then
- languages.rs: 删除从未读取的 LanguageDef.name 字段(语言名即 map key)
- docker.rs: start_container if let Err → ?;timeout match → is_err();
  嵌套 if 合并(修 task 2 遗留 + clippy 1.96 新 lint)
- pages/admin/posts.rs: use_signal 冗余闭包(预存 lint,顺带修绿)
- codemirror editor.ts: biome format 重排 setSchema 单行
2026-07-03 18:57:25 +08:00
xfy
5fdc88a48c fix(runner): resolve task 2 container runner review findings 2026-07-03 18:02:56 +08:00
xfy
d8d30df80c feat(infra): implement Docker execution layer using bollard client 2026-07-03 17:59:52 +08:00
xfy
2fba43dc3f chore(infra): add bollard dependency and modular structure 2026-07-03 17:49:08 +08:00