outrider v1.7Live on the GitHub Marketplace
// for teams improving AI systems

Know your
next move

Start with the AI system you want to improve.

Remyx finds and implements changes worth testing, then measures them against the signals your team already trusts. Each result improves the next recommendation.

works with your stack today github · linear · mlflow · w&b · slack
remyxai-cli · studio.remyx.ai
// seen across the AI community
cerebral_valleymlops_communitypytorch_confodscai_quality_conf
// 00 the_problem

There is no shortage of things to try

Research, issues, roadmaps, production failures, team proposals, and the codebase itself can all surface possible changes. Most are not worth the engineering time. Teams still need to know which changes fit the system well enough to test, and whether those changes improve the metrics they care about.

which change next?
worth the work?
already tested?
yet another paper
another proposal
works on my evals
still no lift
signal or noise?
what should we try now?
// 01 recommendations

The next change worth testing

Remyx checks candidate changes against your codebase, goals, constraints, and past results, then surfaces the strongest candidate and why it fits. If nothing clears the bar, it says so.

  • matched to a live call site in your code
  • checked for fit, reachability, and license
  • across prompts, retrieval, tools, and routing

Remyx gives your coding agents a better place to start and your evaluation stack a concrete change to measure. Your team decides what ships.

nearest win day-to-day iteration high-impact change the change Remyx surfaces

# Remyx aims past the nearest win · illustrative

candidate changes
25
prompts · retrieval · routing
fits your codebase
6
matched to a live call site
3
license checks out
2
high enough confidence
1
1 reviewable artifactor a recorded skip

# from possible changes to one worth testing · illustrative · counts vary by repo and run

// 02 see_it

What happened on real repos

Outrider has opened draft PRs on well-known public AI repos after checking that each change fits the codebase. Two contributions have merged into Hugging Face PEFT after testing, maintainer review, and revision. Other contributions are still under upstream review.

2 merged upstream · 2 under upstream review · skips on most runs by design

huggingface/peft ★ 21k MERGED UPSTREAM · AUG 3, 2026

Riemannian-preconditioned LoRA optimizer

Implemented, tested, benchmarked, refined through maintainer review, and merged upstream, with AI assistance disclosed and a human approving every step.

"this PR is already quite mature"peft maintainer · first review

+401 / −3 · 6 files · 5 commits · 9 tests · 33 days open to merge
huggingface/peft ★ 21k MERGED UPSTREAM · AUG 21, 2026

Super-Tuning & Supra · frozen-weight adaptation

Coordinated with the maintainer before implementation, revised through two review rounds, and merged with a maintainer-requested fix to PEFT's shared tuner utilities included.

The paper's first author co-authored commits on the branchcoordination issue #3450 · maintainer-confirmed before implementation

+1365 / −6 · 28 files · 10 commits · 16 days open to merge
huggingface/lerobot ★ 26k UPSTREAM REVIEW · OPENED JUL 15, 2026

EcotReasoningModule · dense chain-of-thought supervision

Adapts the annotation pipeline to emit embodied chain-of-thought reasoning traces, wired into LeRobot's existing module contract. First-round review comments were addressed the same day.

AI assistance disclosed per LeRobot's AI_POLICY.mdtested end-to-end with an open VLM via an OpenAI-compatible endpoint · no cloud API key required

2 commits · 11 files · 3 unit tests + an end-to-end dataset run

# fork validations demonstrate the pipeline on real codebases. upstream statuses reflect maintainer decisions we don't control. skips happen on most runs by design.

// 03 how_it_works

From a proposed change to evidence

A change can start from your team's own proposal or from Remyx. Remyx implements it in a testable form, measures the result, and records what happened so the next recommendation starts with more evidence. Hover a step to trace it through the cycle.

every step uses real production outcomes start implement measure
EXP-0412 · retrieval-reranker# hover a step
startfrom your team or from Remyx: rerank after retrieval may raise groundedness, matched to your context-build call site
implementdraft PR #214 rerank top-20 → top-5 before context build
measuregroundedness +4.8 pts · answer relevance +2.9 pts · p95 +41ms, in budget
decideteam chooses SHIP · result informs the next recommendation
// 04 validation

Earlier results inform the next recommendation

Remyx keeps evals, benchmarks, CI runs, A/B tests, and production outcomes attached to the change. Later recommendations can use that history as context.

  • your evals, benchmarks, CI, A/B tests, and production signals
  • results stay attached to the change and decision
  • earlier outcomes provide context for later recommendations
a draft PR entersyour policy decides
quality gateskip low-signal PRs
your eval suiteoffline + A/B, your metrics
verdictpass · warn · fail
on passpromote → ready + reviewers
every validated resultsharpens the next

# you set the policy. starts in observe-only.

// 05 own_your_evals

Own your evaluation system

Define quality using the criteria your team already trusts.

Use the evals, benchmarks, and metrics your team already trusts. Remyx runs them consistently against proposed changes and keeps the results and decision history in your repo. Remyx handles orchestration and reproducible execution.

Explore Validate

// 06 integrations

Works with your stack

The tools you already use, tied back to the same change and result. More ship every month.

plan & ship

# planned, shipped, reviewed

githublinearjiraslack+ more

measure & learn

# offline + online results

mlflowwandbarizelangfusestatsiglaunchdarkly+ more

build & run

# implemented + executed

claude-codemodalhuggingface+ more

# bring your own provider key: Anthropic, Z.ai GLM, Moonshot Kimi, and more.

// 07 the_pilot

Prove it on one AI system

Choose one AI product or subsystem and one primary metric. Remyx can start from changes your team already wants to try or surface new candidates, then implement the ones that fit and measure them against your criteria.

the_scope

One repository, one technical owner, one AI product or subsystem, and one primary metric. Success criteria agreed in writing before the pilot starts.

the_guarantee

A scoped install, evaluation criteria reviewed by your team and committed to your repo, an agreed set of changes run through the full workflow, and a final evidence report. Refundable if the guaranteed work is not delivered.

the_terms

Runs in your GitHub Actions by default, with the option to run entirely in your infrastructure. Direct technical support is included. The pilot fee credits toward your first annual agreement.

Request a pilot

// 08 trust

Security that fits how you already work

Remyx provisions server-side through a scoped GitHub App. Access is per repo and revocable, your keys stay in your own repo secrets, and a human gates every merge.

scoped per-repo access keys in your own repo secrets review mode by default human gates every merge SSO & audit logs VPC or self-hosted

Read our security practices →

// 09 who_its_for

For teams building AI systems

Remyx helps AI teams decide what to try next using evidence from their own system. Changes, results, and decisions stay connected so the next recommendation can build on what the team already learned.

$ whoami → ai_engineer

Evaluate changes faster

Remyx keeps the repo context and results from earlier work available when new changes are evaluated.

  • recommendations informed by prior outcomes
  • context on every change
  • faster validation of new ideas
$ whoami → team_lead

Make better decisions

Remyx keeps changes, results, and decisions together so teams can prioritize work using evidence from their own system.

  • visibility across changes and outcomes
  • evidence behind every decision
  • a shared history of what the team tried
// 10 the_team

Built by practitioners, for teams shipping AI in production

Mathematicians and award-winning ML practitioners, a decade applying AI in robotics, healthcare, recommendation, and enterprise data.

Salma Mayorquin

Salma Mayorquin

ceo & co-founder

Applied Math - UC Berkeley. Former Databricks Solutions Architect, startups to Fortune 500. Recognized by NVIDIA's developer community.

Terry Rodriguez

Terry Rodriguez

cto & co-founder

Mathematics - UC Berkeley, UNC Chapel Hill. 10+ years of production ML at Riot Games, Tubi, and Robust.AI. Open-source tools cited by Google DeepMind.

Ready to decide with evidence?

Start free with Outrider to see what it recommends for your repo, or request a pilot on one AI system and one primary metric.

# your next move, with evidence.