I'm curing the behavior of the overlapping tokenization context using a series of experimental mechanisms introduced into AlephLM.
TLDR; The AlephLM components are currently being specialized towards distillery components.
Each component and deviation is specifically chosen to attenuate towards distillation assistance, so these models will likely have other troubles if finetunes are performed on them. They are not official releases, but there are many distillations being created currently.
Most of which are happening considerably faster on my local 4090 card, as they do not require a humongous amount of compute to complete. As I condense the process I will keep the system updated to make sure all major changes, all major trains, and all major offsets to the core AlephLM are recorded.
The upcoming experiments are highly experimental and may both result in something that does not function, while potentially yielding something that trains substantially faster and learns substantially more effectively.
This run is primarily focused on the splat attention. The current mechanism is approaching SDPA in terms of recall but still not there. The generalization of this delta isn't perfect yet, but it's improving.
The distillery however, is showing that full AlephLM models that aren't perfect to SDPA recall, are yielding improved accuracy over SDPA for distillation. This does not mean these mechanisms are capable of guaranteed recall yet, but it does mean that the accuracy currently is capable of distilling information.
Without the splat hubs in Mini-Beatrix, the model cannot function. They are not arbitrary mechanisms, as the disable tests each major validation show the model requires them at every state. Beatrix has gone through growing pains, and the arms are substantially more difficult to analyze than arms on other models, primarily because there are so many more statistics. That being said, the results are substantially pointing to one direction and that direction is unanimously towards a risk mechanism.
The V2 points towards another mechanism. There's no way around it, the yield is strong, the results are strong, but the parity is too close to the current mechanism and the results.
With that, one of the experiments is an atlas lookup trained from the expert itself and then utilized as a codebook reference for the student. These are likely not going to yield, but it's an option to test. Codebooks aren't as strong in practice as the experiments would show, but they are definitely strong.
Two core potential shifts on the model may yield some small changes, but nothing will yield just sitting here idling by with the research.
I have a series of upcoming experiments that will be much more likely to fail. I should say, they need to fail, otherwise progress won't happen.

