jrs

Java Review Swarm · multi-perspective review for Spring & the JVM

The Java reviewer that tries to prove itself wrong.

Eight specialist reviewers read your diff in parallel. Then every finding is cross-examined against the real code by a skeptic whose only job is to refute it. The ones that don't survive never reach you.

Free for non-commercial use · Commercial teams: free 31-day evaluation · Runs inside Claude Code, your code never leaves your environment.

Eight specialist reviewers. Every finding cross-examined against the real code. The bug doesn't survive.

Why it's different

Most AI reviewers hand you a wall of maybes. This one hands you a verdict.

Noise is the reason review bots get muted. JRS is built to earn the opposite reflex — when it flags something, it has already tried to talk itself out of it.

01 / DEPTH

Java-native, not language-agnostic

It knows JPA proxy self-invocation, EAGER-by-default to-ones, expand/contract migration safety, virtual-thread pinning, and @Transactional scope — the pitfalls a generic linter has never heard of.

02 / RIGOR

Verified findings only

Every gating finding is refutation-tested against the real code before you see it. Agreement between reviewers raises priority but never counts as proof — same-model reviewers hallucinate together.

03 / CLARITY

A gate, not a pile

One call over only the findings that survived: 🚫 Blocked · ⚠️ Request changes · 💬 Approve with concerns · ✅ Looks good. Design and style opinions inform — they never block a merge.

Field test — real PRs from Apache, Netty & Halo

We ran it on other people's code, and scored it against the maintainers.

Seven unseen pull requests — DolphinScheduler, Halo, Dubbo, ShardingSphere, plus large ones from Kafka, Pulsar, Flink, and Druid — then eight hundred eighty-seven mined regression pairs: PRs that introduced a real bug the maintainers later fixed, replayed against the introducing PR. The world chose those bugs, not us.

628 / 938
maintainer-fixed regressions detected at the PR that introduced them — real Kafka, Pulsar, Flink, Hadoop, Lucene, Elasticsearch, Cassandra, Trino, Presto, Parquet, Keycloak, Solr, Avro, ORC, netty, Iceberg & more production code
0
false positives across five clean merged PRs — all ✅ LOOKS GOOD, reproduced on a second run
0
fabricated findings across all eight hundred eighty-seven regression-pair runs on real production code

Finder model: Opus 4.8 n=1–70 44/70 (63%); Grok 4.5 n=71–890 537/817 (66%); Opus 5 n=891–941 47/51 (92%) (n=51 — read it as a small new stratum, not a rate); total 628/938 (67%). See the results table Model column.

8 ×
specialist reviewers + adversarial verification, ~2M tokens on a deep run

Read the caveats, they're the point. N is large enough to be informative, but this is still an existence / tripwire result, not a guaranteed rate — the sample is adversarially selected toward weak lanes and harder surfaces; borderline findings were scored as false positives on purpose. The honest pattern behind the 628/938: detection tracks how directly the bug sits on the changed code. On-surface, pointable regressions — a wrong comparison, a missing release, an off-by-one that throws, a downgraded auth check, a self-deadlocking executor — are caught reliably, and so are subtle ones when the owning reviewer reads the out-of-diff helper it depends on (a config path that skips a default, a buffer persisted twice, an executor call that secretly blocks). The misses cluster in emergent cases (a swap that only bites after a failover, an edge case for one id type) and, most tellingly, where the bug's wrongness lives in a different file the diff never touches — and several misses still surfaced a different real bug in the same PR. That last class is now measured rather than asserted: a mechanical screen of both diffs across every pair finds 28% of the misses are not clean in-diff tests (against 16% of the hits), and on the pairs where the defect is on the shown surface recall is 486/707 = 68.7%. Only 3 pairs were provably invalid — one fix branched before its intro merged, one bank named the same PR as both intro and fix, one fix merged seven weeks before its intro — and those were removed from both sides, taking the corpus 890 → 887 and the headline 65.4% → 65.5%. Nothing else was deleted. Full split: recall by attribution stratum. Across all eight hundred eighty-seven runs: zero fabricated findings. A few caught regressions were initially under-rated in severity — each drove a shipped calibration fix. On several pairs it also surfaced defects deeper than the maintainers' shipped fix — code-verified but not maintainer-adjudicated, so we keep them out of the headline. The field-test methodology and every per-PR verdict are public: see the complete field-test results → The regression-pair run artifacts ship in the repo — ask for access and audit them offline.

How it works

Fan out. Fold. Cross-examine. Gate.

A pipeline, not a single prompt — so a bad finding has to get past a skeptic, and no lone agent can swing the verdict.

01 · FAN OUT

Eight independent specialists

Correctness, concurrency, security, persistence, API & ops, performance — each reviews the diff in parallel, and none knows the others exist. Convergence means something only when it's uncoordinated.

02 · FOLD

Merge the duplicates

When several reviewers land on the same defect, it's folded into one finding that keeps every angle — so a single fix-site gates once, not five times.

03 · CROSS-EXAMINE

A skeptic tries to refute it

Each finding faces a verifier that hunts for the guard, the config, the caller-side check that makes it a false alarm. Killing or minting a verdict-deciding call takes a 2-of-3 panel.

04 · GATE

One verdict, only the survivors

A single gated call over the confirmed findings the diff actually introduced. Everything unverified is labeled as such and kept out of the gate — never silently dropped.

Built for the reviewer

Made for Java teams that take review seriously.

Your code stays put

Runs inside Claude Code, in your own environment — no third-party service, no new data processor to get through security. The difference between a git clone and a procurement cycle.

Report-first, never trigger-happy

The default deliverable is a report with paste-ready comments. It posts to a PR only when you ask, always as a comment — never an automated approve or block. Humans decide merges.

It compounds with your outages

Feed it your team's postmortems and house conventions; they become failure patterns and repo-specific rules the reviewers apply on every future run.

Three tools, one brain

A full swarm review, a quick single-specialist inline pass, and a delta re-review that checks whether an author addressed your last round — sharing one set of Java reviewer briefs.

Licensing

Free to evaluate. Fair to run a business on.

Source-available under PolyForm Noncommercial — access by request, then read it, run it, audit every prompt it sends. Commercial use is a separate license.

Non-commercial
Free
  • Full capability — all eight reviewers and verification
  • Personal projects, research, education, evaluation
  • Use by non-profit & government organizations
  • Source access by request · PolyForm Noncommercial 1.0.0
Request trial access
Commercial
From $300 / team / month, billed annually
  • Free 31-day evaluation first — no signature, no card
  • Commercial-use license for for-profit work
  • Reviewing company code, in your engineering process
  • Priority on new failure patterns & Java-version support
  • Optional onboarding of your own postmortems
Get a commercial license

Design partners wanted. The first few teams that run JRS on their real pull requests get six months of commercial use free, in exchange for a short monthly report — what it flagged wrongly, what it missed, whether the verdicts earned trust. Your postmortems become failure patterns; your false positives make the verifier stricter. Become a design partner →