Research & evidence

Built on evidence, not demos

HyperC sits between a research lab and a product company: working papers, a proposed data contract for the category, a defined falsification protocol — and a commercial product measured in customer-realized profit. The research program is bigger than one model: HyperC is the research platform for robots that make money, and the laboratory is the market itself. This page keeps the evidence inspectable, with synthetic and live results clearly separated.

How to read this page. Synthetic stress test results come from controlled benchmark markets with known ground truth — they demonstrate mechanism, not live profitability. Live deployment results are company-reported production outcomes — note: not audited by a human licensed auditor unless stated. Historical results do not guarantee future outcomes.
Notebook results

The selection-bias trap, measured

From the executed slower-market-waves notebook in the Computable Markets working paper (August 2026). The point is not that P34 wins — it's why conventional models lose.

Synthetic stress test
0.9266 AUC → −$417.4k
The conventional baseline's great backtest, then reality

A tuned XGBoost classifier scored 0.9266 ROC-AUC on the business-observed holdout and showed +$61.7k in naive observed-data evaluation — then lost $417.4k applied to the full 20-week future opportunity menu it had never been forced to refuse.

Synthetic stress test
+$2,250 on $4,146
P34, executed REST-API run on the withheld week-80 menu

P34 selected 115 trades, predicted +$2,404, realized +$2,250. The calibrated XGBoost baseline selected 177 trades and realized −$8,789 on $18,821 deployed. Live system invocation on a synthetic market — not a live-market trial.

Synthetic stress test
+$14.8k vs −$219.8k
Difficult benchmark: P34 vs tuned conventional baseline

Partial observability, selection bias, regime change, optimistic backtest traps. In the easy stationary/growing benchmark P34 preserved upside: +$228.6k vs the baseline's +$227.1k.

Synthetic stress test
99.4%
of evaluated orders rejected in the published study

Disciplined refusal is the central mechanism: no-trade is a rewarded output, and false-positive control is explicit in the model class.

Production

Live deployments

$30M+
sales generated for customers, >95% of trades unsupervised
Company-reported, cumulative since 2023
~$100M/yr
reseller operated with limited supervision
Thousands of signals; POs and shipping orders generated
3,000+
loans issued by the model in a live lending test
Micro-credit vertical, 2025

Company-reported figures. Note: not audited by a human licensed auditor. Detailed per-market performance packs (evaluated counts, rejection rates, predicted-vs-realized economics, win/loss distributions) are published per reference deployment as part of the evidence engine.

Publications

Papers, benchmarks & code

Working paper · Aug 2026

Computable Markets: Business Menus, Sales Event Tapes, and Cross-Market Profit-Directed Learning

The operator-relative theory of computable markets, the Menu-Sales-Description (MSD) data contract, the Computable-Market Thesis, application templates for 10+ verticals, and a falsification protocol. Includes the revised notebook results above.

PDF Download the working paper
Technical report · Jun 2026

P34: Learning When Not to Trade

The P34/PARML technical report: benchmark construction (partial observability, selection bias, regime change), baseline tuning, results and limitations — with notebooks. Presented at computablemarkets.com.

GIT hyperc-ai/p34-technical-report — notebooks & results
Patent

Menu-structured trade selection with many-worlds model ensembles

Provisional patent filed on the P34 architecture: synthetically-trained universe selection, Pareto-front alignment, and the confidence interface. Multiple defensive publications published.

Protocol

Evaluation & falsification protocol

How to try to break P34: holdout menu construction, baseline calibration rules, refusal accounting, and what a falsifying result would look like. Ask your own AI to audit it.

PDF Evidence pack — drop file to publish
The research agenda

Making money is only the first question

A robot that can pursue profit creates a much larger set of questions. How should it manage risk? What should it refuse to do? How much autonomy should it have? How should it account for consequences that never show up in its P&L? How should many such systems cooperate — or compete? HyperC treats these as engineering and economic research questions to be studied alongside profitability, not after autonomous business is deployed at scale.

Profit & decision quality

Can the system create positive economic outcomes from biased, incomplete, changing real-world data?

Risk & survival

Can it avoid strategies that look profitable until they destroy the business?

Alignment & control

Can owners specify objectives, constraints, approvals and stop conditions that stay meaningful as autonomy grows?

Fairness & distribution

Who captures the efficiency these systems create — and how should they behave toward customers, suppliers, workers and competitors?

Sustainability & externalities

What happens when a profitable action imposes costs that never enter the robot's objective?

Cooperation

Can autonomous businesses share data, coordinate capacity and create more value together than apart?

Market structure

When many robots learn and compete at once, do margins compress? Does capital concentrate? What new structures emerge?

Governance

Who can inspect, constrain, pause, audit or change the objective of a system that participates in the economy continuously?

None of these are solved. They are part of why the platform exists — the working stack turns them from philosophy into measurable experiments. Our laboratory is the market.

Limitations

What we tell investors and customers up front

P34 is for computable markets; it does not yet automatically prove that a market is computable.

If the deployment-time menu is far from training context, an explicit refusal path is not yet exposed.

The business must have had more options historically than it chose — otherwise there is insufficient policy-selection signal.

Latency ranges from under a minute to hours; very high-frequency applications are outside current scope.

Generalization across all markets is not fully known; pilots proceed through shadow testing and controlled deployment.

Unknown unknowns remain — P34 is an early release in a new model class.