Program Projects Publications FDE About Contact 中文
RRLAB — INDEPENDENT AI RESEARCH LAB

Let multiple AIs collaborate
on what no single AI can solve.

RRLab is Rodney Rui's independent research lab.
Research question: how can multi-model systems be reliable
within known boundaries, creative beyond them, and self-governing?
Research findings become open-source tools;
tools power engineering delivery.

202/202
Subtasks completed
96 : 76
Ablation result
25.9%
Reasoning depth gain
13
Zenodo preprints
01

Research Program
Reliable × Creative × Self-Governing

A single model has limits. RRLab does not search for a stronger
model — it asks how to organize multiple models, each with its
own strengths, into a system that delivers reliably within known
boundaries, produces verifiable new knowledge beyond them, and
generates its own governance rules.
Four research lines, each with published evidence.

01

Dimension-Direct Routing

Explicitly decompose what capability a task needs, then route
directly to the strongest model along that dimension, without
intermediate representations. Reasoning depth improves by 25.9%
over implicit routing.

→ Paper: Dimension-Direct Routing(DOI: 10.5281/zenodo.19504212

02

Uncertainty Routing (BPL/OA)

A purely classical framework inspired by the measurement
structure of quantum mechanics: before a decision, the system
retains a distribution over plausible worlds, instead of
collapsing prematurely into a single confidence score. That
distribution maps to an auditable route: auto-accept, review,
fallback, abstain. The goal is not to prevent AI from being
wrong — it is to prevent AI from being wrong with high confidence.

→ Paper: BPL/OA: Auditable Uncertainty Routing(DOI: 10.5281/zenodo.19936937

03

Resonance Exploration

Hallucination cannot be dammed, but it can be channeled. In
exploration mode, multiple models independently analyze the
same object from complementary perspectives — not seeking
consensus, but detecting resonance: structurally compatible
interpretations reached from different paths. Resonant
directions enter a computational falsification queue. In the
BSD campaign, 6/7 strongly resonant hypotheses survived
subsequent verification.

→ Paper: Deliberation and Exploration(DOI: 10.5281/zenodo.19505278
→ Case study: From Resonance to Rigor(DOI: 10.5281/zenodo.21500421

04

AI Governing AI

Orchestration rules are not stacked from human presets — a
reasoning model derives them deductively from system invariants
alone. Ablation: deductive 96/100, inductive 76/100. In a
larger experiment, a system given only minimal constraints
(explore + detect resonance) self-organized dependency-aware
batching and cross-model file sharing — the paper calls it
"ordered behavior under minimal constraints": observed, not
explained. Consistent cross-project experience: adding human
micro-constraints degraded output quality.

→ Paper: LLM-Skill Orchestration(DOI: 10.5281/zenodo.19503674

02

Projects
Where the research lands

01

Multi-Model Orchestration SystemIn development · 5 papers published

The runtime that carries the four research lines. Multi-model
orchestration with a multi-dimensional evaluation layer
(50-task benchmark), supporting both deterministic and
exploratory modes. Largest experiment to date: 15 skills,
7+ heterogeneous models, 77 tool calls, 74 minutes of
autonomous operation.

Link: github.com/rrlab-tech (opening up alongside RRLab-Bench)

02

eVoiceClaw Product FamilyFour product lines in parallel development

Full-stack productization of multi-model orchestration.
eVoiceClaw (iOS voice assistant, multi-provider switching),
Desktop (multi-agent workstation), Pro (fully local knowledge
base for doctors and engineers), Cerebellum (on-device
"cerebellum" model: semantic routing + privacy awareness +
safety review). Built end-to-end: Swift frontend, Python
backend, agent orchestration.

03

RRLab-Benchv0.3.3 released · Open source

Agent code-modification capability audit. It does not rank
models; it answers two engineering questions: after a code
change, do previously passing tests still pass (FRR)? At equal
quality, who is faster and cheaper?

Link: github.com/rrlab-tech/rrlab-bench

04

Voice ToolkitTwo layers, first layer released

Layer 1 — Mac-TTS (v0.1 released, open source): turns every
Mac's built-in speech into a Unix pipe. Zero disk I/O, pure
CPU inference, TTS + STT in one. Layer 2 (in development):
emotion-controlled multi-character voice acting, currently
via cloud TTS.

Link: github.com/rrlab-tech/mac-tts

05

Bench-HarnessResearch in progress

From "benchmarking models" to "benchmarking and evolving the
harness": how the runtime around a model — workflows,
evaluation, permissions — can itself be systematically
evaluated and improved.

03

Publications
The program, in writing

All preprints are on Zenodo; every DOI is verifiable.
ORCID: 0009-0004-3898-4655

Track 1: AI Multi-Model Orchestration

validated on the self-built multi-model orchestration system

Track 2: Computational Number Theory / BSD Conjecture

application of the resonance-exploration mode; large-scale SageMath computation + Magma verification + AI-assisted research

04

FDE Consulting
Research capability, delivered per project

Independent Forward Deployed Engineer.
RRLab's research is not paper theory — the four orchestration
modes, benchmark evaluation, and uncertainty routing are
engineering methods that can be embedded directly into your
systems. I work embedded in your team and take LLMs from demo
to production-line productivity.

202/202
All subtasks passed in a 50-task agentic benchmark
96 : 76
Ablation: deductive rule generation beats human preset stacking
13
Preprints, methodology fully public and auditable
05

About Rodney Rui

Mathematician × AI systems builder × Independent FDE.

Academic background in computational number theory: numerical verification methods for elliptic curves and the BSD conjecture. Recent focus: AI systems engineering — multi-model orchestration, agent evaluation, uncertainty routing. RRLab is the unified vehicle for this work.

This methodology was not read from reports. In the BSD campaign, a multi-model AI system produced publishable discoveries — resonance proposed hypotheses, deterministic computation falsified them, and 13 papers were the byproduct. Built the eVoiceClaw product family from scratch across four product lines. Self-built benchmarks quantify model capabilities; every selection decision has a measurement behind it.

Frontier exploration on genuinely open problems, with results — most business problems are easier than this.

06

Contact

Have an AI idea and don't know where to start? Let's talk — no pressure.