Start with the AI system you want to improve.
Remyx finds and implements changes worth testing, then measures them against the signals your team already trusts. Each result improves the next recommendation.
Research, issues, roadmaps, production failures, team proposals, and the codebase itself can all surface possible changes. Most are not worth the engineering time. Teams still need to know which changes fit the system well enough to test, and whether those changes improve the metrics they care about.
Remyx checks candidate changes against your codebase, goals, constraints, and past results, then surfaces the strongest candidate and why it fits. If nothing clears the bar, it says so.
Remyx gives your coding agents a better place to start and your evaluation stack a concrete change to measure. Your team decides what ships.
# Remyx aims past the nearest win · illustrative
# from possible changes to one worth testing · illustrative · counts vary by repo and run
Outrider has opened draft PRs on well-known public AI repos after checking that each change fits the codebase. Two contributions have merged into Hugging Face PEFT after testing, maintainer review, and revision. Other contributions are still under upstream review.
2 merged upstream · 2 under upstream review · skips on most runs by design
Implemented, tested, benchmarked, refined through maintainer review, and merged upstream, with AI assistance disclosed and a human approving every step.
"this PR is already quite mature"peft maintainer · first review
+401 / −3 · 6 files · 5 commits · 9 tests · 33 days open to mergeCoordinated with the maintainer before implementation, revised through two review rounds, and merged with a maintainer-requested fix to PEFT's shared tuner utilities included.
The paper's first author co-authored commits on the branchcoordination issue #3450 · maintainer-confirmed before implementation
+1365 / −6 · 28 files · 10 commits · 16 days open to mergeAdapts the annotation pipeline to emit embodied chain-of-thought reasoning traces, wired into LeRobot's existing module contract. First-round review comments were addressed the same day.
AI assistance disclosed per LeRobot's AI_POLICY.mdtested end-to-end with an open VLM via an OpenAI-compatible endpoint · no cloud API key required
2 commits · 11 files · 3 unit tests + an end-to-end dataset runSteers RL fine-tuning toward the tokens the model is least sure about, so each update lands where it teaches the most.
The Outrider run that started the chain. Validated on our fork, then coordinated and reshaped into huggingface/peft #3382.
Turns repeated user corrections into a runtime check, so the agent stops making the mistakes you already corrected.
Swaps the hard clip in RL fine-tuning for a smooth penalty that gives rare tokens more room to move.
Lets a robot policy learn from imperfect demonstrations instead of throwing the messy data away.
# fork validations demonstrate the pipeline on real codebases. upstream statuses reflect maintainer decisions we don't control. skips happen on most runs by design.
A change can start from your team's own proposal or from Remyx. Remyx implements it in a testable form, measures the result, and records what happened so the next recommendation starts with more evidence. Hover a step to trace it through the cycle.
Remyx keeps evals, benchmarks, CI runs, A/B tests, and production outcomes attached to the change. Later recommendations can use that history as context.
# you set the policy. starts in observe-only.
Define quality using the criteria your team already trusts.
Use the evals, benchmarks, and metrics your team already trusts. Remyx runs them consistently against proposed changes and keeps the results and decision history in your repo. Remyx handles orchestration and reproducible execution.
The tools you already use, tied back to the same change and result. More ship every month.
# planned, shipped, reviewed
# offline + online results
# implemented + executed
# bring your own provider key: Anthropic, Z.ai GLM, Moonshot Kimi, and more.
Choose one AI product or subsystem and one primary metric. Remyx can start from changes your team already wants to try or surface new candidates, then implement the ones that fit and measure them against your criteria.
One repository, one technical owner, one AI product or subsystem, and one primary metric. Success criteria agreed in writing before the pilot starts.
A scoped install, evaluation criteria reviewed by your team and committed to your repo, an agreed set of changes run through the full workflow, and a final evidence report. Refundable if the guaranteed work is not delivered.
Runs in your GitHub Actions by default, with the option to run entirely in your infrastructure. Direct technical support is included. The pilot fee credits toward your first annual agreement.
Remyx provisions server-side through a scoped GitHub App. Access is per repo and revocable, your keys stay in your own repo secrets, and a human gates every merge.
Remyx helps AI teams decide what to try next using evidence from their own system. Changes, results, and decisions stay connected so the next recommendation can build on what the team already learned.
Remyx keeps the repo context and results from earlier work available when new changes are evaluated.
Remyx keeps changes, results, and decisions together so teams can prioritize work using evidence from their own system.
Mathematicians and award-winning ML practitioners, a decade applying AI in robotics, healthcare, recommendation, and enterprise data.
ceo & co-founder
Applied Math - UC Berkeley. Former Databricks Solutions Architect, startups to Fortune 500. Recognized by NVIDIA's developer community.
cto & co-founder
Mathematics - UC Berkeley, UNC Chapel Hill. 10+ years of production ML at Riot Games, Tubi, and Robust.AI. Open-source tools cited by Google DeepMind.
Start free with Outrider to see what it recommends for your repo, or request a pilot on one AI system and one primary metric.
# your next move, with evidence.