Skip to content
View reacher-z's full-sized avatar

Highlights

  • Pro

Block or report reacher-z

Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
reacher-z/README.md

Hi, I'm Yuxuan Zhang

PhD student at University of British Columbia · Affiliated with Vector Institute · Research in AI Agents, LLMs, and RL

Website Research Program X Google Scholar


EMNLP 2026


  Top Projects

ClawBench: Can AI Agents Complete Everyday Online Tasks? GitHub Stars

EMNLP 2026 Findings · V1: 153 tasks · V2: 130 tasks · 144 live websites · 15 categories

Paper · Dashboard · Dataset · PyPI

VidGround: Watch Before You Answer GitHub Stars

Visually grounded post-training for video LLMs.

Paper · Project Page · HF Paper

Dr. Claw: Your AI Research Assistant GitHub Stars

EMNLP 2026 System Demonstrations · A full-stack research workspace for taking projects from idea to paper.

Homepage · npm · Releases

OpenSkill: Open-World Self-Evolution for LLM Agents GitHub Stars

EMNLP 2026 · Builds both skills and verification signals from scratch, without target-task supervision.

Website · Paper · HF Paper

RewardHarness: Self-Evolving Agentic Post-Training GitHub Stars

COLM 2026 · A self-evolving agentic reward framework for image-editing evaluation.

Project Page · Paper · HF Paper · Releases


  Research Guides & Runnable Examples

Research question Try the teaching resource Original work
Does valid structured output preserve the requested fields and values? StructEval CPU mini-lab · Two-minute walkthrough StructEval
Does a retrieved paper actually support the claim being written? Citation evidence worksheet ScholarCopilot
What does a text-only video-question probe establish? VidGround probe-log mini-lab Watch Before You Answer

These coauthor-maintained resources use handwritten examples, not new model experiments. The citation worksheet is a separate reading exercise, not a ScholarCopilot capability. Retaining a question after a text-only probe does not by itself prove visual dependence. Scripts, inputs, scope notes and original sources are linked from each resource; the video uses synthetic narration.

All publications and citation materials · Research program


  GitHub Activity

GitHub Contribution Graph


  News


  Contact

 yuxuan.zhang(at)ubc.ca      Google Scholar      GitHub      X      Website

Pinned Loading

  1. TIGER-AI-Lab/ClawBench TIGER-AI-Lab/ClawBench Public

    Open-source benchmark for browser AI agents on daily tasks.

    Python 905 62

  2. vidground vidground Public

    Watch Before You Answer: Learning from Visually Grounded Post-Training (arXiv 2604.05117)

    Python 65

  3. HarnessBench HarnessBench Public

    Benchmark for comparing agent harnesses on everyday online tasks — fixes the base model, varies the harness. Sister project of ClawBench, same scoring pipeline.

    Python 56 4

  4. GraphEngineering GraphEngineering Public

    Vendor-neutral graph orchestration for reliable agentic systems: one portable Graph IR, native TypeScript/Python runtimes, deterministic barriers/routers, and cross-language conformance.

    TypeScript 5 1

  5. TIGER-AI-Lab/RewardHarness TIGER-AI-Lab/RewardHarness Public

    Self-evolving agentic reward framework for image-editing evaluation — 47.4% on EditReward-Bench from only 100 preference demos, no reward-model training. arXiv 2605.08703.

    Python 63 9