Graph
Map assumptions, evidence and conclusions.
GERO Research
Original research and independent analysis of the problems frontier laboratories are naming — translated into mathematical questions, assurance mechanisms and tests that can fail.
The name
GERO organizes independent checks around AI-generated claims and keeps every conclusion inside the boundary of its evidence.
Map assumptions, evidence and conclusions.
Apply bounded checks to necessary claims.
Expose contradictions and unresolved risk.
Coordinate models, tools and expert review.
The journal
We do not copy laboratory articles. We link to the primary source, state what it reports, separate our inference and propose a falsifiable GERO response.
83 publications
Publication records across GitHub, GERO, LinkedIn and Zenodo →
A finite zero-rate coefficient is formed as 0/0, and a nearby put can exceed the boundary solver's iteration limit. Released-wheel reproduction, analytic limit, independent CRR control and a green upstream correction.
Three single- and double-digital paths use the wrong probability or discount term. Released-wheel reproduction, payoff-derived controls, falsification and a green upstream correction.
A documented two-date call can dereference the absent reference-period end when an annual interval reaches February 29. Released-wheel reproduction, independent 91/366 control and submitted correction.
Released PyXIRR 0.10.8 returns an unconverged vector RATE and fails at zero starting guesses. Independent discounting, compiled Rust checks and a tested candidate correction.
A list of distinct expiries is reduced to its final horizon and broadcast across every option. Released-wheel reproduction, independent oracle and submitted correction.
A component-axis alignment error turns parallel vectors into a nonzero cross product. Clean native CPU regression and a bounded candidate correction.
Clean native CPU evidence of integer arithmetic before float promotion. Tested bounded repair, unchanged inexact controls, and unresolved large-value range cases.
A zero rate raises ZeroDivisionError; tiny nonzero rates can leave a material unpaid or overpaid balance in synthetic amortization checks.
Dirty price is scaled by the requested face while accrued interest remains at face one, violating the clean-value identity in a released-wheel reproduction.
Finite means become infinity in both the reference evaluator and Runtime CPU provider because the input-type sum overflows before division.
Finite Euclidean norms become infinity or zero in both the reference evaluator and Runtime CPU provider because values are squared before scaling.
A public Fineract case with 25 common checks, scoped candidate corrections, reproducible sources and a 40-second customer overview. Synthetic component evidence; customer impact is not established.
Finite L1 and L2 inputs collapse to zero in both the reference evaluator and Runtime CPU kernel because avoidable intermediate norms overflow or underflow.
Clean CPU reproduction of a small-matrix slogdet range defect. Tested FP32 workaround; LU autodiff limitation and rejected repair counterexample retained.
Finite rows at 1e30, 1e-30, 1e300 and 1e-300 produce NaN, zero, unnormalized values, or the wrong dtype in the released reference and Runtime CPU implementations.
Clean native CPU reproduction: an absent weight narrows finite float32 bias. A one-line candidate changes 106 failed assertions out of 385 to zero; 16 controls remain unchanged.
Standard GroupNorm returns zeros for finite FP16 inputs because a sum of squares overflows before division. A CPU research patch preserves channel grouping: 45 main-suite failures become zero, with 41 compatibility checks passing.
A finite float16 batch makes the default running variance infinite, leaving later evaluation at zero. A CPU research patch changes 102 main-suite failures to zero, with 33 compatibility checks passing. Existing corrupted state is not restored.
Float16 [3072,4096] should clip to approximately [0.6,0.8], but MLX returns zero. A CPU research patch passes 1701 assertions across 219 scenarios, plus 39 compatibility and autodiff checks.
For C/(C*t), the derivative at t=1 must remain −1. MLX returns zero or negative infinity at extreme scales. A research patch passes 1328 native comparisons, plus 1886 compatibility checks.
For atan2(C*t,C), the derivative at t=1 must remain 0.5. MLX returns zero or infinity at extreme scales. A research patch passes 1368 native comparisons; both failed intermediate patches are preserved.
At float16 x=1000, arcsinh and arccosh return finite values but zero derivatives. A normalized real-derivative patch passes 518 native comparisons across four dtypes; GPU and speed are untested.
At x=-20, reverse mode returns zero for a derivative near 0.20611536. A local VJP repair plus the known exp control passes 216 native comparisons; the VJP-only patch can regress float64.
At two equal float32 inputs of 100000000, MLX returns gradient [1,1] instead of [0.5,0.5]. An isolated CPU study separates a new VJP prototype from an already reported exp limitation.
A smooth x²/2 example returns second derivative 0 instead of 1. A research C++ prototype passes 477 checks; GPU and large-tensor performance remain untested.
An extra output-mask factor suppresses its derivative at zero. A one-line C++ repair passes 377 scenarios and 604 checks; the demonstrated training step lowers loss from 18 to 0.
With transpose=False, affine gather_qmm miscomputes scale and bias gradients. A local C++ orientation repair passes 205 scenarios and 379 checks; the measured quadratic update then lowers the loss.
Valid sorted-index calls produce incorrect gather_mm and affine gather_qmm gradients. A bounded local FP32 CPU repair passes 69 scenarios and 626 checks; quantized FP16 fallback remains a limitation.
MLX uses the forward Hadamard transform in reverse differentiation for nonsymmetric factors 20 and 28. A local C++ repair passes 64 scenarios and 624 checks; measured energy steps confirm the correction.
MLX power returns NaN derivatives at zero for smooth fixed-exponent polynomials. A local C++ patch passes 41 scenarios and 396 checks; prior work and test-discovery limits are documented.
MLX mapped boolean assignment routes gradients across batch examples. A local C++ patch passes 19 scenarios and 255 checks; explicit loops and actual-forward finite differences agree.
A strict C++ scatter_max/min block update loses a winning gradient. Independent finite differences confirm the reference; a minimal extent correction passes 26 targeted checks.
nntrainer returns correct forward values but loses three quarters of the input-gradient sum in a minimal asymmetric-padding case.
A native C++ nntrainer case study: zero-rate Dropout forwards correctly but misroutes multiple input gradients. Finite differences, a local patch and 19 passing tests.
Native MLX 0.32.2 CPU reproduction: finite Gaussian NLL with a variance gradient of the wrong sign. Inputs, references, regression tests and limits of a mitigation.
A native CPU reproduction isolates a shape-dependent Muon update in MLX. Equivalent convolution and linear layers agree before the optimizer step; a minimal reorder passes 31 focused tests.
A real loss increases after a gradient-descent step in MLX. Native tests isolate missing complex conjugation and an arccosh branch error; the local C++ patch passes 131 scenarios.
MLX returns NaN second derivatives for a smooth cumulative-product polynomial at zero. A division-free C++ prototype passes 57 local scenarios, with an explicit O(N log N) cost.
Slices [inf, -inf] and [-inf, -inf] both become NaN, although the exact results and ONNX Runtime give +inf and -inf respectively.
Inference uses blended training statistics and five-output training cannot return its requested statistics because mode is selected from momentum instead of output count.
With sorted=0 and a nonzero axis, first-occurrence indices are always applied to axis 0, producing wrong data or an immediate out-of-bounds exception.
Inputs of equal rank with smaller non-gather index dimensions are mathematically well-defined and accepted by ONNX Runtime, but the reference oracle requires exact shape equality.
Reproduced cosine NaNs in MLX and an incorrect LayerNorm input gradient in nntrainer, with tested local patches and downloadable evidence.
OpenAI reports rapidly growing agent use while warning that activity metrics are hard to interpret; the mathematical task is to estimate validated progress under steering, selection and changing compute.
Anthropic's Fermat formalization shows the scale now possible in Lean; the remaining assurance problem is preserving meaning, provenance and dependency scope from source theorem to checked root.
A reproducible nntrainer causal-mask report: future values affect earlier outputs. Native tests, a local repair and an upstream issue are available.
The input [1, -1] becomes [0, 0], while the exact L1-normalized result and ONNX Runtime both return [0.5, -0.5].
The summaries of versions 18 and 21 contradict the num_groups attribute and the operator's own four-channel, two-group example.
Reproduced BCE and logaddexp failures in MLX and nntrainer, with mathematical references, tested local patches and downloadable evidence.
Reproduced ELU, GELU and Softplus failures in pinned MLX and nntrainer code, with tested local repairs and downloadable evidence.
A 104-logit gap makes the reference loss infinite while the mathematical result and ONNX Runtime output remain finite.
All inputs remain JavaScript safe integers, yet binary64 rounding changes a 3-unit allocation from the exact [2, 1, 0] to [3, 0, 0].
Scalar base 240 is accepted and formatted as 1.10 for the exact value 250/240, while the equivalent multi-base currency [20, 12] is correctly rejected.
Exact and Wilson bounds, sample-size thresholds, runnable code and the assumptions that make a zero-error headline meaningful or misleading.
123 graphs, 246 evaluations, seven flagged graphs. A repeated CPU benchmark, fourteen archive replays and an 80-digit diagnostic check, with downloadable evidence and explicit limits.
An omitted reduction axis throws during initialization; a batch-two regression exposes the broken default and verifies a 23-test local correction.
A two-value counterexample shows the CPU float16 path returning [0, 0] where the ONNX rounding contract requires [1, -1].
An accepted 7-quadrillion-unit amount makes two exact residuals collapse to zero in binary64, assigning the final unit to the wrong share.
Fineract's public number-of-payments helper returns 7 when its adjacent payment function was given a 12-period loan.
A five-line counterexample shows reciprocal multiplication changing a result that direct complex division preserves, with a separate integer-power phase-error finding.
A large-negative-real fast path flips component signs in complex cosh and sinh, propagates into cos and sin, and is reinforced by incorrect test expectations.
Five bounded audit cases show how documented equivalence, exact identities, input preservation and parameter propagation expose failures that fixed expected values can miss.
MLX 0.32.2 returns a valid affine parameterization that differs from the formula stated in its public documentation.
Five compact case studies show how chain rules, shape semantics, special values and execution parity expose silent failures in production AI software.
A one-scalar experiment shows how an in-place label negation makes repeated KL-divergence backward calls alternate their gradient sign.
Four of five implemented loss derivatives omit the configured loss scale; four runtime counterexamples and a 71-test repair expose the cross-component contract failure.
An author preprint proving fractional projective-gap inequalities in the complete second chaos, all split-product degrees and the first fully mixed harmonic cubic case.
An exact arithmetic check found that one Swift Numerics example reported 3 where its stated rounding rule and its own explanation require 2.
Double.root(3125, 5) returns 5.000000000000001 although the exact, representable result is 5.0.
A three-value UINT8 input collapses a negative endpoint and real zero into the same code. The pinned evidence package records a candidate ratio-formula repair and 15 passing local tests.
A research perspective connecting the reliability problems named by frontier labs to observability, control, verification and calibrated abstention.
A falsification-first protocol for turning AI-generated mathematical arguments into dependency graphs, bounded verdicts and replayable evidence.
Why scope, evidence, calibration, falsification, escalation and traceability should be a single contract around every high-consequence AI answer.
Krol proved the leading narrow-cone constant in 1973 and left the next term as O(1). Identifying it took analysis; establishing who owned the first term took archive retrieval.
Six claims, zero arithmetic errors, five failures. Every one failed above the arithmetic layer: wrong quantity, prior art, or a fabricated citation.
The dataset, the command and the full result behind the 60-pair claim — including the distinction between a refuted defect and an abstention that the short phrasing hides.
Anthropic's faithfulness results suggest that a visible chain of thought is a noisy sensor, not a guaranteed transcript of computation.
Mechanistic interpretability can expose computational structure, but the map is partial and the meaning assigned to it remains a scientific claim.
OpenAI's correction to a First Proof submission exposes false acceptance as the central metric for AI-assisted mathematics.
OpenAI's monitorability evaluations raise a systems question: will a useful safety signal remain informative after training, scaling and distribution shift?
Google DeepMind's Aletheia separates research mathematics from bounded contests and makes revision and failure admission part of the system.
Google Research separates finding a mistake from repairing it; GERO turns that distinction into separate evidence contracts.
DeepMind's IMO result is a major capability milestone; the next mathematical question is how performance transfers beyond a bounded contest distribution.
Evidence standard
Every frontier brief preserves those four layers. Read the editorial and evidence standard or follow new work through RSS.