This incident exposes the main problem of the frontier model market. Kimi K3, with its 2.8 trillion parameters and 1-million token context window, requires colossal amounts of VRAM for inference. Amid a chip shortage, the startup pragmatically chose to protect SLAs for existing clients over chasing new ones. The way out of the infrastructure bottleneck will be the release of full open weights, announced for July 27. By releasing the architecture to open-source, Moonshot AI relieves itself of the inference burden, forcing corporations to run the heavyweight agent on their own (on-premise) servers.
Source: Moonshot AI / Reuters
LLMMoonshot AIInfrastructureB2BChina