Aolani, a Singapore-founded neocloud, has launched the Aolani Token Factory, a managed inference platform that enables organizations to deploy and scale AI models on a pay-per-token basis without the need to provision or manage underlying GPU infrastructure. This launch makes Aolani the first Singapore-founded neocloud to offer production-grade, managed inference at scale, a significant step in addressing the growing demand for AI infrastructure in the region.
As global AI companies expand operations in Singapore and businesses worldwide invest in AI to drive tangible outcomes, the need for robust inference infrastructure has accelerated. The Aolani Token Factory is designed to bridge the accessibility gap, providing AI-native companies and enterprises with a compliant, high-performance path from experimentation to production-scale deployment. By offering per-token metering, customers can pre-purchase credits and pay based on token consumption, avoiding the capital-intensive investment typically associated with GPU infrastructure. Aolani manages the entire inference stack—including GPU capacity allocation, model serving, orchestration, scheduling, and workload optimization—allowing customers to scale usage without continuously provisioning additional hardware.
The platform supports leading open-source models at launch, including DeepSeek, GLM, Kimi, and Qwen, with plans to expand the catalogue based on customer demand. Customers can also deploy their own models through OpenAI-compatible APIs. For enterprises with strict compliance and data residency requirements, dedicated capacity and data isolation options are available. The Token Factory is tailored for three core production use cases: AI agents for high-volume inference and workflow automation, enterprise AI applications such as internal copilots and knowledge assistants, and coding agents for code generation, completion, testing, and review.
Sea Xu, Applied AI Research Lead at Aolani, emphasized the platform's technical strengths: "The Aolani Token Factory is built on a high-performance inference stack that supports the most in-demand open-source model families. We designed the platform for fast model adaptation and deployment, so our customers can get access quickly as new models emerge. As Southeast Asia's AI ecosystem evolves and grows rapidly, it is our goal to ensure that the infrastructure serving it keeps pace."
Nicholas Chia, Chief Executive Officer at Aolani, highlighted the business impact: "Fast-moving AI natives want to build and ship products flexibly and on-demand, without the need to manage GPU fleets. With our competitive per-token pricing and a fully managed stack, companies can go from model selection to production deployment without the capital outlay or operational complexity of self-managed infrastructure. This is a significant milestone for us and our customers as the Aolani Token Factory will fundamentally change how customers access AI compute."
The launch comes at a time when AI adoption in Asia is surging, and the Token Factory aims to lower the barriers to entry for organizations of all sizes. By offering a flexible, cost-effective model, Aolani is positioning itself as a key enabler of AI innovation in the region. Interested parties can register interest at Aolani Token Factory.

