Java Review Swarm · multi-perspective review for Spring & the JVM
The Java reviewer that tries to prove itself wrong.
Eight specialist reviewers read your diff in parallel. Then every finding is cross-examined against the real code by a skeptic whose only job is to refute it. The ones that don't survive never reach you.
Free for non-commercial use · Commercial teams: free 31-day evaluation · Runs inside Claude Code, your code never leaves your environment.
Eight specialist reviewers. Every finding cross-examined against the real code. The bug doesn't survive.
Why it's different
Most AI reviewers hand you a wall of maybes. This one hands you a verdict.
Noise is the reason review bots get muted. JRS is built to earn the opposite reflex — when it flags something, it has already tried to talk itself out of it.
Java-native, not language-agnostic
It knows JPA proxy self-invocation, EAGER-by-default to-ones, expand/contract migration safety, virtual-thread pinning, and @Transactional scope — the pitfalls a generic linter has never heard of.
Verified findings only
Every gating finding is refutation-tested against the real code before you see it. Agreement between reviewers raises priority but never counts as proof — same-model reviewers hallucinate together.
A gate, not a pile
One call over only the findings that survived: 🚫 Blocked · ⚠️ Request changes · 💬 Approve with concerns · ✅ Looks good. Design and style opinions inform — they never block a merge.
Field test — real PRs from Apache, Netty & Halo
We ran it on other people's code, and scored it against the maintainers.
Seven unseen pull requests — DolphinScheduler, Halo, Dubbo, ShardingSphere, plus large ones from Kafka, Pulsar, Flink, and Druid — then eight hundred eighty-seven mined regression pairs: PRs that introduced a real bug the maintainers later fixed, replayed against the introducing PR. The world chose those bugs, not us.
Finder model: Opus 4.8 n=1–70 44/70 (63%); Grok 4.5 n=71–890 537/817 (66%); Opus 5 n=891–941 47/51 (92%) (n=51 — read it as a small new stratum, not a rate); total 628/938 (67%). See the results table Model column.
Read the caveats, they're the point. N is large enough to be informative, but this is still an existence / tripwire result, not a guaranteed rate — the sample is adversarially selected toward weak lanes and harder surfaces; borderline findings were scored as false positives on purpose. The honest pattern behind the 628/938: detection tracks how directly the bug sits on the changed code. On-surface, pointable regressions — a wrong comparison, a missing release, an off-by-one that throws, a downgraded auth check, a self-deadlocking executor — are caught reliably, and so are subtle ones when the owning reviewer reads the out-of-diff helper it depends on (a config path that skips a default, a buffer persisted twice, an executor call that secretly blocks). The misses cluster in emergent cases (a swap that only bites after a failover, an edge case for one id type) and, most tellingly, where the bug's wrongness lives in a different file the diff never touches — and several misses still surfaced a different real bug in the same PR. That last class is now measured rather than asserted: a mechanical screen of both diffs across every pair finds 28% of the misses are not clean in-diff tests (against 16% of the hits), and on the pairs where the defect is on the shown surface recall is 486/707 = 68.7%. Only 3 pairs were provably invalid — one fix branched before its intro merged, one bank named the same PR as both intro and fix, one fix merged seven weeks before its intro — and those were removed from both sides, taking the corpus 890 → 887 and the headline 65.4% → 65.5%. Nothing else was deleted. Full split: recall by attribution stratum. Across all eight hundred eighty-seven runs: zero fabricated findings. A few caught regressions were initially under-rated in severity — each drove a shipped calibration fix. On several pairs it also surfaced defects deeper than the maintainers' shipped fix — code-verified but not maintainer-adjudicated, so we keep them out of the headline. The field-test methodology and every per-PR verdict are public: see the complete field-test results → The regression-pair run artifacts ship in the repo — ask for access and audit them offline.
How it works
Fan out. Fold. Cross-examine. Gate.
A pipeline, not a single prompt — so a bad finding has to get past a skeptic, and no lone agent can swing the verdict.
Eight independent specialists
Correctness, concurrency, security, persistence, API & ops, performance — each reviews the diff in parallel, and none knows the others exist. Convergence means something only when it's uncoordinated.
Merge the duplicates
When several reviewers land on the same defect, it's folded into one finding that keeps every angle — so a single fix-site gates once, not five times.
A skeptic tries to refute it
Each finding faces a verifier that hunts for the guard, the config, the caller-side check that makes it a false alarm. Killing or minting a verdict-deciding call takes a 2-of-3 panel.
One verdict, only the survivors
A single gated call over the confirmed findings the diff actually introduced. Everything unverified is labeled as such and kept out of the gate — never silently dropped.
Built for the reviewer
Made for Java teams that take review seriously.
Your code stays put
Runs inside Claude Code, in your own environment — no third-party service, no new data processor to get through security. The difference between a git clone and a procurement cycle.
Report-first, never trigger-happy
The default deliverable is a report with paste-ready comments. It posts to a PR only when you ask, always as a comment — never an automated approve or block. Humans decide merges.
It compounds with your outages
Feed it your team's postmortems and house conventions; they become failure patterns and repo-specific rules the reviewers apply on every future run.
Three tools, one brain
A full swarm review, a quick single-specialist inline pass, and a delta re-review that checks whether an author addressed your last round — sharing one set of Java reviewer briefs.
Licensing
Free to evaluate. Fair to run a business on.
Source-available under PolyForm Noncommercial — access by request, then read it, run it, audit every prompt it sends. Commercial use is a separate license.
- Full capability — all eight reviewers and verification
- Personal projects, research, education, evaluation
- Use by non-profit & government organizations
- Source access by request · PolyForm Noncommercial 1.0.0
- Free 31-day evaluation first — no signature, no card
- Commercial-use license for for-profit work
- Reviewing company code, in your engineering process
- Priority on new failure patterns & Java-version support
- Optional onboarding of your own postmortems
Design partners wanted. The first few teams that run JRS on their real pull requests get six months of commercial use free, in exchange for a short monthly report — what it flagged wrongly, what it missed, whether the verdicts earned trust. Your postmortems become failure patterns; your false positives make the verifier stricter. Become a design partner →