Glassbox Agent Harness
Glassbox-Agent-Harness面向AI Agent 的飞行记录器与评测实验室,让每一次行动可观察,每一份上下文可追溯,每一次优化都有证据。
SELECTED WORK
面向AI Agent 的飞行记录器与评测实验室,让每一次行动可观察,每一份上下文可追溯,每一次优化都有证据。
Agent Arena — 让 Agent 在真实任务中比赛、被裁判审查、留下证据、积累声誉的智能体竞技场。Real-task evaluation harness for LLM agents with judge review and reputation.
AI-native software engineering harness — 18 agents, 9 workflows, one closed loop: Issue → Worktree → Plan → Build → Review → Evidence → Merge → Memory.
A low-token visual evidence compiler for text-only coding agents. Convert images into compact Visual Evidence Packets (VEP) for DeepSeek, Codex, Claude Code, and other text-only models.
Autonomous multi-agent economic sandbox — A2A protocol agents compete, trade, and govern inside a simulated economy with a God's Eye observability dashboard.
Personal second-brain dashboard — Obsidian vault sync, knowledge graph, daily notes, and task tracking in one local-first UI.
多模态人体状态感知 → 中医体质分析 → 智能按摩方案推荐 | OPLRI + TCM + EEG pretrained models · Gradio demo
AMD Hackathon contribution: an AI-native live-event poster engine that combines music rhythm, stage lighting, and real-time visual composition.
An interactive Chinese source-code guide to LLM inference: 13 browser experiments make continuous batching, KV cache, paged attention and scheduling visible.
A personal trading operations dashboard that brings market data, positions, backtests and execution checks into one TypeScript workspace.
An image-to-collectible skill that turns a supplied subject into a faithful, commercially believable packaged object.
On-chain agent intent verification for a Monad hackathon: deterministic evidence checks followed by a fail-closed human signing gate.