Artifacts
and Open Tools

AI Frontier
We traced the absolute frontier of AI performance and how well can LLMs perform tasks in perfect conditions.
Thesean AI
A Lab Building Best Execution for LLMs. Reducing LLM Costs 50% Using Best-Execution for Intelligence.
K-Steering
Python package supporting Llama, Gemma, and Qwen. Control Multiple Behaviors in Language Models at Once
Ares
RL-first framework for training coding agents with true online reinforcement learning.
Code Review Bench
An Unbiased OSS Benchmark For Code Review Agents Measure which tools actually catch bugs and improve code.