Grok 4 Fails to Deliver in Alpha Arena Benchmark Despite Aggressive Positioning
Grok 4, a large language model developed by SpaceXAI, participated in the Alpha Arena benchmark between October 17 and November 3, 2025. The model traded perpetual contracts on the Hyperliquid venue with five other models, including DeepSeek V3.1 and Gemini 2.5 Pro. Grok finished at -39 percent on the final leaderboard, while DeepSeek ended at +48 percent.
The benchmark measured directional conviction over a seventeen-day window across six correlated assets, not durable edge. The model's initial aggressive positioning paid off during the first few days but produced a significant loss when the market regime shifted.
SpaceXAI documentation states that Grok models have no knowledge of current events beyond their training data unless external search tools are explicitly enabled for individual requests. This means that without enabling Web Search or X Search, every price and listing detail returned by the model is a reconstruction from training data.