What just i read in model card.

#14
by shivshankar - opened

Is it some kind of frekenstine merge of totally different base model

There are two options: either this is a breakthrough that deserves a peer-reviewed paper, or it's total bullshit. For now, I'm sticking with the second one.
I know for sure that transferring knowledge between different architectures by merging their blocks without retraining is complete nonsense. Unless I misunderstood the process.

Yes this is a cross-architecture project. That's not really unheard of, there was a reddit post about doing it with Wan and LTX. You couldn't do this with LLM's, at least I don't think so. But if anything credit and more attention needs to go to the Minimax H3 design and architecture for supporting it as a real possibility and this is just proof of concept. If future models were similar there might be more cross-architecture functionality. Or this is a quirk that unbiased deterministic models excel at being poked and prodded.
A model like LTX is much worse for it. I did the reverse and sent H3 back into the LTX the same way; it definitely helped slightly in ways that it could with a little bit of a cleaner prompting output and response, but it was slight since expanding the scope breaks LTX very quickly.

There are two options: either this is a breakthrough that deserves a peer-reviewed paper, or it's total bullshit. For now, I'm sticking with the second one.
I know for sure that transferring knowledge between different architectures by merging their blocks without retraining is complete nonsense. Unless I misunderstood the process.

Couldn't agree more. This is clearly a joke. It's wild that so many people actually believe this garbage. And seeing veteran YouTubers hype it up? Absolutely ridiculous.

I have 3 extracted loras and three checkpoints put out that are training-less using the same method. Can all be A/B'd and look at the changes that each one brings v.s. base. Despite that still got a bunch of idiots posting without looking or trying it.

There are two options: either this is a breakthrough that deserves a peer-reviewed paper, or it's total bullshit. For now, I'm sticking with the second one.
I know for sure that transferring knowledge between different architectures by merging their blocks without retraining is complete nonsense. Unless I misunderstood the process.

Couldn't agree more. This is clearly a joke. It's wild that so many people actually believe this garbage. And seeing veteran YouTubers hype it up? Absolutely ridiculous.

https://huggingface.co/TenStrip/10Eros-Max/blob/main/h3_graft_methodology.md#machine-learning-case

There are two options: either this is a breakthrough that deserves a peer-reviewed paper, or it's total bullshit. For now, I'm sticking with the second one.
I know for sure that transferring knowledge between different architectures by merging their blocks without retraining is complete nonsense. Unless I misunderstood the process.

Couldn't agree more. This is clearly a joke. It's wild that so many people actually believe this garbage. And seeing veteran YouTubers hype it up? Absolutely ridiculous.

I have no dog in this fight whatsoever, but anyone who used Sulphur and Eros for LTX will definitely confirm that whatever they did there worked, and worked well. So, all things considered, I'm inclined to believe TenStrip. He's been very careful to say that this is highly experimental and that he's just trying different techniques to see what works. I've never once heard him claim this is as some sort of breakthrough; on the contrary, he's repeatedly said this has been attempted in the past, but that he just wants to see how it works with H3 since the model's structure seems more receptive to it.

Don't really get the hate. Let the man tinker, even if it ultimately doesn't result in anything substantial. It literally costs you nothing, and yet you still feel the need to be complete shitbags about it. Maybe get out of the basement, evaluate your own life, and try making something of it.

I can upload the full strength one and you can all use a really bad forced 16fps version of H3 that's over-influenced by Wan with fully divergent motions and prompt response to prove my point.

TenStrip changed discussion status to closed

There are two options: either this is a breakthrough that deserves a peer-reviewed paper, or it's total bullshit. For now, I'm sticking with the second one.
I know for sure that transferring knowledge between different architectures by merging their blocks without retraining is complete nonsense. Unless I misunderstood the process.

Couldn't agree more. This is clearly a joke. It's wild that so many people actually believe this garbage. And seeing veteran YouTubers hype it up? Absolutely ridiculous.

I have no dog in this fight whatsoever, but anyone who used Sulphur and Eros for LTX will definitely confirm that whatever they did there worked, and worked well. So, all things considered, I'm inclined to believe TenStrip. He's been very careful to say that this is highly experimental and that he's just trying different techniques to see what works. I've never once heard him claim this is as some sort of breakthrough; on the contrary, he's repeatedly said this has been attempted in the past, but that he just wants to see how it works with H3 since the model's structure seems more receptive to it.

Don't really get the hate. Let the man tinker, even if it ultimately doesn't result in anything substantial. It literally costs you nothing, and yet you still feel the need to be complete shitbags about it. Maybe get out of the basement, evaluate your own life, and try making something of it.

I'm not here to fight. I don't have enough time for that. I was just leaving my thoughts here so anyone passing by takes these claims with a grain of salt, because it's frustrating to see non-technical people spreading this nonsense all around as fact. Playing with statistical random noise and getting a biased result isn't knowledge transfer. Do you understand what transferring knowledge between different architectures without training actually implies? It means no more model training ever. Does that honestly sound reasonable to you?

I don't care what he does in his own time, as long as he keeps it to himself. There’s a reason researchers follow strict methodologies and only reveal discoveries after extensive verification.

Edit : I just see he changed is md file. Letting AI spread nonsense on your own repo seems like a weird move, and it proves he really doesn't understand what he's doing. Anyway, this discussion has been closed.

image

There are two options: either this is a breakthrough that deserves a peer-reviewed paper, or it's total bullshit. For now, I'm sticking with the second one.
I know for sure that transferring knowledge between different architectures by merging their blocks without retraining is complete nonsense. Unless I misunderstood the process.

Couldn't agree more. This is clearly a joke. It's wild that so many people actually believe this garbage. And seeing veteran YouTubers hype it up? Absolutely ridiculous.

I have no dog in this fight whatsoever, but anyone who used Sulphur and Eros for LTX will definitely confirm that whatever they did there worked, and worked well. So, all things considered, I'm inclined to believe TenStrip. He's been very careful to say that this is highly experimental and that he's just trying different techniques to see what works. I've never once heard him claim this is as some sort of breakthrough; on the contrary, he's repeatedly said this has been attempted in the past, but that he just wants to see how it works with H3 since the model's structure seems more receptive to it.

Don't really get the hate. Let the man tinker, even if it ultimately doesn't result in anything substantial. It literally costs you nothing, and yet you still feel the need to be complete shitbags about it. Maybe get out of the basement, evaluate your own life, and try making something of it.

I'm not here to fight. I don't have enough time for that. I was just leaving my thoughts here so anyone passing by takes these claims with a grain of salt, because it's frustrating to see non-technical people spreading this nonsense all around as fact. Playing with statistical random noise and getting a biased result isn't knowledge transfer. Do you understand what transferring knowledge between different architectures without training actually implies? It means no more model training ever. Does that honestly sound reasonable to you?

I don't care what he does in his own time, as long as he keeps it to himself. There’s a reason researchers follow strict methodologies and only reveal discoveries after extensive verification.

Edit : I just see he changed is md file. Letting AI spread nonsense on your own repo seems like a weird move, and it proves he really doesn't understand what he's doing. Anyway, this discussion has been closed.

You're beyond a clown. It's not block transfer it's guided attention shift of the H3 model, and works on other transformers too. Cross-model attention influence needs to be studied more and there's nothing wrong with having more universality in open source. Get the fuck out of the community. Blocked.

The interesting question isn't whether cross-architecture weights can magically transfer “knowledge.” They obviously don't in the conventional sense. The interesting question is whether carefully mapped, magnitude-preserving perturbations of semantically corresponding subspaces can induce reproducible donor-specific behavior in the target model. If the effect survives norm-matched random controls, donor/block permutation tests, and systematic Q/K/V/MLP ablations, then dismissing it as “random noise” is no longer a sufficient explanation. At that point, we can argue about what to call the phenomenon—but the phenomenon itself would be real.

真正值得探讨的问题,并非跨架构权重能否神奇地迁移“知识”——显然,在传统意义上它们并不能做到这一点。真正有趣的问题在于:针对语义对应子空间,若施加经过精心映射且保持幅值(magnitude-preserving)的扰动,能否在目标模型中诱导出可复现的、源自特定“捐赠者”(donor)的行为特征?如果这一效应在经过范数匹配的随机对照、捐赠者/模块置换检验以及针对 Q/K/V/MLP 的系统性消融实验后依然存在,那么将其简单归结为“随机噪声”便不再足以解释该现象。届时,我们可以探讨该如何命名这一现象,但现象本身将是确凿无疑的。

The interesting question isn't whether cross-architecture weights can magically transfer “knowledge.” They obviously don't in the conventional sense. The interesting question is whether carefully mapped, magnitude-preserving perturbations of semantically corresponding subspaces can induce reproducible donor-specific behavior in the target model. If the effect survives norm-matched random controls, donor/block permutation tests, and systematic Q/K/V/MLP ablations, then dismissing it as “random noise” is no longer a sufficient explanation. At that point, we can argue about what to call the phenomenon—but the phenomenon itself would be real.

The key to doing more with it is expanding it to a training process instead of just blending calculations. Training one model's full attention on to the other would more cleanly transfer it. Data is not relevant, the only main thing carried is next frame behavior and prompt response as the model learns to attend to prompt and flow differently. Right now it's only useful for H3 because it's hard to train but it's also more robust and unified and can survive the grafting better. I did the reverse process of H3 into LTX2.3 attention and there was the same overall effect, but higher strength breaks the fragile LTX output and you can't even touch it's attn_2 layer tokenization without breaking LTX. It'd be much more useful on models like Zimage or Krea2, image models where you also don't need to worry about audio or next-frame motion flow quality as well.

Sign up or log in to comment