SparkX 2.5: A 4B Model With 1M-Token Context
SparkX 2.5 pairs a 4-billion-parameter model with a configured context window of just over one million tokens, using a hybrid attention layout to keep most of that history cheap. The architecture cuts the KV cache to roughly a quarter of what full attention across every layer would