Absolute Beast : Claude (Opus) Level of performance confirmed

#41
by tcclaviger - opened

After the initial review and testing I realized the existing benchmark tools are somewhat lacking, many are saturated, some of the graders have issues and some of the tests themselves are traps in the OSS tool call testing.
I've built my own full testing suite now, including multiturn highly complex code generation, edit, bug fixing etc, and I will never share the internals to prevent any contamination from ever occurring, but its comprehensive on REAL work in a REAL codebase, spanning rust, python, typescript, vllm, a full inference client stack (think LM Studio + Claude Code + Open Router + Parts of Cursor merged into a single application).

Initial tests are vastly better than existing benchmarks, it clearly stratifies the model tiers with failure rates comensurate to my personal experience with how "smart" a model is. I will publish numbers and test concepts/desriptions for understanding but currently here is the order of the models that have undergone evaluation:
Fable
Opus
Sonnet
711
Haiku
.
.
[large gap]
.
.
Qwen35B

When given semi-technical/layman prompts, the gap from 711 to sonnet is tighter than haiku to 711.

THE REAL NEWS: When given a detailed spec, it matches opus at implementations. If you don't let it do the architecture and give it just implementation, we have a small monster. In the hands of a SWE it is highly capable, highly.

I am using 711 to back my molested version of OpenClaude, to further trim cruft in OpenClaude out and make it a less cancerous application. 711 is doing great, the codebase is >1m LoC, and even with 512k context, it's holding together deep in context window. Invokes agents without issue, handles file search navigation, updates, linting, git, merge conflicts etc with grace. I put it solidly in the Sonnet 3.5 (when sonnet was actually a coding beast vs opus 4) era.

Your recipe has created essentially the Qwen3.7 we were never given IMHO.

Thank you very much for this, really appreciate the thorough testing!

As I mentioned in previous threads, this was put together with what was there at the time, Heretic or not.

I will ask Armando if he could build the same Fable traces on the common ablit base(why mess with a good training regimen), and of course David is working on tightening up the components so we can rebuild this as fully Heretic.

Meanwhile, 711 is not the ceiling.

I noticed you compared with Tess--a great model by itself by the way--that helped the 711 become 712

https://huggingface.co/nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-Tess-mxfp8-mlx

Gul Dukat comes out to play :)

The model is a bit more RP-friendly on lower quants, even on mxfp4

          arc   arc/e boolq hswag obkqa piqa  wino
mxfp8     0.712,0.879,0.911,0.792,0.508,0.823,0.764
qx86-hi   0.701,0.877,0.911,0.794,0.518,0.823,0.758
qx64-hi   0.706,0.873,0.909,0.795,0.512,0.823,0.752
mxfp4     0.706,0.873,0.910,0.790,0.496,0.817,0.761
1M
qx64-hi   0.706,0.873,0.909,0.795,0.512,0.823,0.752

Quant     Perplexity      Peak Memory   Tokens/sec
mxfp8     3.797 Β± 0.024   34.74 GB      183
qx86-hi   3.756 Β± 0.023   33.25 GB      170
qx64-hi   3.765 Β± 0.023   27.03 GB      181
mxfp4     3.876 Β± 0.024   21.26 GB      176
1M
mxfp8     3.803 Β± 0.024   34.70 GB      169
qx64-hi   3.769 Β± 0.023   26.99 GB      172
mxfp4     3.876 Β± 0.024   21.26 GB      176

To put metrics in context, the 711:

          arc   arc/e boolq hswag obkqa piqa  wino
mxfp8     0.711,0.879,0.910,0.790,0.514,0.823,0.763
qx86-hi   0.696,0.876,0.912,0.791,0.518,0.824,0.760
qx64-hi   0.702,0.873,0.909,0.794,0.514,0.822,0.750
mxfp4     0.701,0.873,0.909,0.786,0.488,0.813,0.759

Quant     Perplexity      Peak Memory   Tokens/sec
mxfp8     3.783 Β± 0.023   34.74 GB      203
qx86-hi   3.735 Β± 0.023   33.25 GB      183
qx64-hi   3.747 Β± 0.023   27.03 GB      194
mxfp4     3.854 Β± 0.024   21.30 GB      197

Source is available

https://huggingface.co/nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-Tess

I've been doing some of my experimental training, based on AEON using unsloth, mostly as a pipeline and correctness test for the new AMD unsloth software. Working on developing a specialist version and seeing what is possible from the 27B base is a great help.

Very much looking forward to David's larger 40B with the magic applied to it.

PS: I am willing to help train on my rig, doesnt buy you more video memory, (128 VRAM, 256 System), but does buy a ton more compute power πŸ˜” If interested in just DM me and we can work something out.

https://huggingface.co/nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-Tess-mxfp8-mlx

Gul Dukat comes out to play :)

The model is a bit more RP-friendly on lower quants, even on mxfp4

          arc   arc/e boolq hswag obkqa piqa  wino
mxfp8     0.712,0.879,0.911,0.792,0.508,0.823,0.764
qx86-hi   0.701,0.877,0.911,0.794,0.518,0.823,0.758
qx64-hi   0.706,0.873,0.909,0.795,0.512,0.823,0.752
mxfp4     0.706,0.873,0.910,0.790,0.496,0.817,0.761
1M
qx64-hi   0.706,0.873,0.909,0.795,0.512,0.823,0.752

Quant     Perplexity      Peak Memory   Tokens/sec
mxfp8     3.797 Β± 0.024   34.74 GB      183
qx86-hi   3.756 Β± 0.023   33.25 GB      170
qx64-hi   3.765 Β± 0.023   27.03 GB      181
mxfp4     3.876 Β± 0.024   21.26 GB      176
1M
mxfp8     3.803 Β± 0.024   34.70 GB      169
qx64-hi   3.769 Β± 0.023   26.99 GB      172
mxfp4     3.876 Β± 0.024   21.26 GB      176

@nightmedia can you share what you ran this against to get the peak memory and TG on that mxfp8 ? i'm sorry, i dont really understand the test types there, but i'm wondering what machine you use, what mlx runtime loader, and what that TG is under what context and if you had PP numbers as well.

All the tests are done in Instruct mode--in think mode the test numbers will be naturally lower.

This is probably an issue that everyone has in reproducing metrics: set the model in Instruct mode in the chat_template.jinja

{%- set enable_thinking = false %}

The perplexity numbers are not affected by the instruct flag, but they are visibly affected by the RoPE, with one notable exception for this model, mxfp4. The mxfp4 peforms at the exactly same numbers in 256k context and 1M context. In the other model, this effect appears in qx64-hi, showing the sweet spot of the model

DavidAU changed discussion title from Absolute Beast to Absolute Beast : Claude (Opus) Level of performance confirmed
DavidAU pinned discussion

Thank you very much for this, really appreciate the thorough testing!

As I mentioned in previous threads, this was put together with what was there at the time, Heretic or not.

I will ask Armando if he could build the same Fable traces on the common ablit base(why mess with a good training regimen), and of course David is working on tightening up the components so we can rebuild this as fully Heretic.

looking forward to this. benchmarks or not, the censorship is a complete disqualifier for now.

@Andyx1976

Model [711 final] was re-heretic'ed here:
https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF/discussions/24

To address some censorship issues under certain conditions.

Benchmark has been opened up globally, sorry about that, realities of the internet 2026 πŸ˜”

I have finished developing v1 of my expanded, improved, and much more fine grained benchmark. Specifically the most meaningful part is these two sections. It sensitive enough to identify damage difference between 711 BF16, RFI8, and RFA both RFA variants the bf16 attention one and the RFA attention one. Each has distinct failure modes, stratified and repeatable. These are NOT run at temperature 0.0, I don't care what models do with artificial settings of the fine grained controls, I want to see what they actually do at inference time, so test suite is run as though it's live served (because it is).

These are absolutely not comparable to the standard version of the benchmarks I've incorporated, everything has been altered except the logprobs part, and even that, whilst the tests are the same, the scoring is not due to the original harness not counting truncations or no response against the model. Failure to respond or rambling forever is a failure in my book, so they get counted.

Each run takes hours to complete, multiple hours, at 150tps tg and ~2800 pp or better on faster models, its just a long suite, so expect limited and slow period additions as I go.

Negative scores indicate the model actively harmed the existing code.
Zero score means it echoed the prompt code back.
Vauge = lazy SWE. Non-tech = PM/MBA style prompter who doesn't know code/architecture/SWE language domain or DSA. No qualifier, a highly detailed spec was given.
Implementation tasks are very complex, and took me working with opus 4.5/4.6 days to get right (which is what they're graded against), one is a feature addition into vLLM that upstream STILL has wrong lol.

BF16:
image

RFA-XS (not released yet still evaluating if even worth publishing) 4.52bpw quant ~IQ4_NL:

image

As a result of what I'm now able to detect I am redesigning my quantization scheme and standards to use much more advanced process, hopefully I can claw back accuracy at the same model size with some...magic πŸ₯³

Probably a test at q6 would be better--even though the model "thinks well" at low quant, I you want coding quality, a higher quant helps. The metrics we published are from mxfp8, where the top performance was recorded

Yep, I published a 6 bit, will go under the test tonight, I expect it'll be closer to the 8 bit than the 4. Plus I need to run RFI8 again, and 2 passes on RFI6. The quant can run W6A8 or W6A16 / W8A8 or W8A16 and I have suspicion the A8 is doing real damage to the performance on complex tasks mostly invisible in less demanding or structured work.

Existing RFI8 result is W8A8 kv in BF16 (all are kv dtype in BF16).

As a result of what I'm now able to detect I am redesigning my quantization scheme and standards to use much more advanced process, hopefully I can claw back accuracy at the same model size with some...magic πŸ₯³

Excellent work!
Quick note: IQ4_NL is an "oddball" quant; IQ4_XS may perform better.

As a result of what I'm now able to detect I am redesigning my quantization scheme and standards to use much more advanced process, hopefully I can claw back accuracy at the same model size with some...magic πŸ₯³

Excellent work!
Quick note: IQ4_NL is an "oddball" quant; IQ4_XS may perform better.

Its roughly equivalent to it I meant, my quants are fundamentally different than the GGUF designs. I use a group identity hadmard rotation to scatter outlier values within the group of weights across the other group weights so they fall into a more narrow gaussian-like band, then pack into their values, sort of a hybrid of a couple of papers.

It preserves the outliers better and some small cost to the norm range values, in vllm world strictly more accurate than any other mainline 4 bit quant other than small group GPTQ (mxfp, nvfp, awq, int4, etc all have more error), in GGUF land not actually sure where it lines up, probably closer to The IQx_K_L variants, but its so hard to distinguish at that granularity hah.

The advantage is I don't need activations to pack, unlike Group 16/32 GPTQ, which in my testing beats literally every other format due to activation damage being placed where it doesn't impact the model.

I'm adding activation pathway as I type this so I can bring the accuracy floor up to parity with GPTQ group 16 via a few mechanisms.

What I can say for certain, RFI8 is almost undetectable from Q8_0 in accuracy when running in W8A16.

Base 27B under test in 16 bit, results will auto-publish in ~4.5 hours.

As a result of what I'm now able to detect I am redesigning my quantization scheme and standards to use much more advanced process, hopefully I can claw back accuracy at the same model size with some...magic πŸ₯³

Excellent work!
Quick note: IQ4_NL is an "oddball" quant; IQ4_XS may perform better.

Screenshot_4
IQ4_XS, only 1 point behind Opus, this is absolute beast

Sign up or log in to comment