Valerois AI Logo
ValeroisAI
Bifrost CSL
Empirical Evaluation Suite

Master Benchmark Scorecard

Standardized empirical evaluation across multi-hop reasoning, code synthesis, and context understanding. Sub-quadratic architectures verified against open-weight baselines.

63× Data & Compute Efficiency Breakthrough

Valkir 16L reached 54.0% HumanEval AST validity on only 4.74 Billion tokens trained on a single AMD RX 9070 XT. Competing models required 300 Billion tokens on multi-node GPU clusters to achieve ~15–20%.

63.3×
Token Ratio
37.2k
Tokens/sec
Valerois SOTA (Glowing)Standard Baselines
Click column headers to sort metrics
Showing 0 of ... models
Model ArchitectureParams
👑 Composite Avg
HumanEval AST
Code Synthesis
AI2 ARC-Easy
Reasoning
WinoGrande
Language
HellaSwag
Commonsense
Tokens
VRAM
Hardware Verification

AMD Radeon RX 9070 XT (16 GB VRAM) running ROCm 7.2 with PyTorch 2.5 Inductor kernel fusion.

Zero-Shot Precision

Strict greedy decode without temperature tuning or prompt engineering tricks. Standard lm-evaluation-harness.

Reproducibility

BF16 checkpoints verified with SHA-256 integrity hashes. Free to inspect and benchmark on consumer hardware.