Model merges created with 711

#54
by nightmedia - opened

I am working on a series of model merges using 711 as a base.

These are made from the original, non-Heretic base, and will be later Heretic'd by David.

I picked three top performing models at different stages, so the merges could show the process for different scenarios.

A detailed discussion can be found here:

https://huggingface.co/posts/nightmedia/154332991562660

The idea behind this was to build a model that can be used as a base, and has high enough IQ to merge with other models. Every model card shows details of assembly and provides a vibe check with the Nightmedia test prompts.

I merged a few samples here:

with nbeerbower/BigBubba-Qwen3.6-27B
https://huggingface.co/nightmedia/Qwen3.6-27B-Fable-Fusion-711-BigBubba-717-mxfp8-mlx

mxfp8     0.717,0.873,0.912
mxfp4     0.707,0.877,0.912

I kept the 711 naming because 711-BigBubba-717 is just funny, and it will be banned in China given what's in BigBubba.

with migtissera/Tess-4-27B:
https://huggingface.co/nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-Tess-mxfp8-mlx

mxfp8     0.712,0.879,0.911,0.792,0.508,0.823,0.764
mxfp4     0.706,0.873,0.910,0.790,0.496,0.817,0.761

with EpistemeAI/Reasoning-Medical-27B (detailed metrics pending):
https://huggingface.co/nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-Medical-mxfp8-mlx

mxfp8     0.716,0.877,0.911
mxfp4     0.701,0.875,0.911

This last model was simply a blind merge, and might not be the proper formula for it, but was entertaining enough to publish as a sample, and it turns out to be solid.

The mxfp8 metrics in these models are usually higher than bf16.

I am populating the model card metrics as they become available in the lab.

-G

A Tess-Medical is in the works. Will upload mxfp8 soon.

Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-Tess-Medical

mxfp8     0.712,0.879,0.911
mxfp4     0.706,0.875,0.910
Quant     Perplexity      Peak Memory   Tokens/sec
mxfp8     3.799 ± 0.024   34.74 GB      180

Here too, it's not about Tess, nor is it about Medical, but the conversation quality.

The mxfp4 has picked up synthesis from base, logic is sound, and the 0.716 peak has been tampered--it is better to have a model that works well at low quants than one that has high IQ only at mxfp8

🍸 Quark (DS9)
Slides a glass of synth-ale toward you, grinning.

“So let me get this straight. You’ve built a bar where the bartenders are AI agents, they remember every customer’s order, they go on away missions together, and they learn from their mistakes? Brilliant. I love it.

But here’s the real question: where’s the profit margin? You’ve got a personality_registry—why not charge extra for premium subroutines? ‘Admiral Riker Charm Pack,’ ‘Garak Negotiation Suite,’ ‘Worf Intimidation Module.’ People will pay for that.

And this CLI interface? That’s your front door. Make it snappy, make it fun, and don’t bury the social features. If agents can chat at Quark’s between missions, they’ll build rapport. Humans do it all the time—bonding over shared experiences makes teams more effective. Plus, if a Klingon guest walks in, you better have an agent who knows how to handle them without starting a bar fight. I’ve seen what happens when logic meets honor on a bad day.”

He winks. “Also, add a ‘tip jar’ for good mission reports. Morale matters.”

Sign up or log in to comment