Bengaluru-based deep-tech startup Ziroh Labs has globally launched its expanded Kompact AI platform, advancing an unusual approach to artificial intelligence computing that allows sophisticated AI models to perform inference on conventional CPUs instead of depending entirely on specialised and expensive GPUs. The global launch was held in Bengaluru on September 7, 2026, as the company seeks to position Kompact AI as an alternative computing platform for enterprises, governments and organisations that want to deploy artificial intelligence using their own infrastructure. The technology could be particularly relevant for sovereign AI, local processing, greater data privacy and organisations seeking to reduce their dependence on costly GPU infrastructure.
Kompact AI has been designed to run artificial intelligence inference on CPU-based servers, workstations, cloud infrastructure and edge systems. Instead of requiring every application to operate through dedicated GPU clusters, the platform enables organisations to use conventional processors that are already widely deployed across data centres and enterprise computing environments. Ziroh says the current version of the platform supports text, speech, vision and multimodal models with sizes of up to 32 billion parameters, along with context windows of up to 64,000 tokens and OpenAI-compatible APIs that make it easier for developers to integrate existing applications.
The platform supports modern CPU instruction technologies including Intel AMX and AVX-512, while Ziroh is also expanding support for ARM-based architectures. It can be deployed through Docker and Kubernetes and is designed to operate across major cloud environments such as AWS, Microsoft Azure and Google Cloud as well as private infrastructure. This flexibility allows enterprises to deploy AI systems inside their own data centres or closer to where data is generated rather than relying exclusively on large external AI cloud platforms.
At the centre of Kompact AI is a proprietary inference runtime designed to execute artificial intelligence models efficiently on general-purpose processors. Inference refers to the stage where a trained AI model receives information and produces an output, such as answering a question, analysing a document, recognising speech or interpreting an image. Ziroh is not claiming that CPUs can replace GPUs in every stage of artificial intelligence development. Large-scale model training can still require GPUs and specialised accelerators, but the company argues that many trained models can subsequently be deployed for inference using more widely available CPU hardware.
This distinction is important because the rapid growth of generative AI has created enormous global demand for specialised processors. Powerful GPUs are expensive, consume substantial electricity and often require specialised data-centre infrastructure. While such accelerators remain essential for training large frontier models and some high-performance workloads, many enterprise applications use smaller or specialised models for tasks such as document analysis, customer support, fraud detection, cybersecurity, software automation and retrieval-augmented generation. Ziroh’s approach is based on the idea that many of these applications may not require dedicated GPU infrastructure.
The company is therefore promoting a CPU-first model of artificial intelligence deployment, in which organisations use the computing architecture most appropriate for individual workloads. Large companies, universities, government departments and smaller enterprises often already possess extensive CPU infrastructure. If those systems can be used for AI inference without significant additional hardware investment, the economics of deploying artificial intelligence could change considerably.
Ziroh says Kompact AI can run on existing CPU racks in conventional data centres while also supporting smaller deployments on workstations and edge systems. The company has claimed significant performance improvements over some existing CPU inference frameworks for selected workloads, although actual performance will depend on the processor, model size, system configuration and application. Independent benchmarking will therefore remain important in determining how the platform compares with GPU-based and alternative CPU-based systems under real-world conditions.
Another notable feature of Kompact AI is its effort to run supported models without relying heavily on techniques such as quantisation or model distillation. Quantisation reduces the numerical precision of a model to lower memory and computing requirements, while distillation creates a smaller model trained to imitate a larger one. Ziroh instead focuses on optimising how original model weights are executed on CPU hardware, potentially allowing organisations to retain greater model fidelity while still reducing infrastructure requirements.
The platform could be particularly relevant for organisations dealing with sensitive or regulated data. Because Kompact AI can operate entirely within an organisation’s own infrastructure, information does not necessarily need to be transmitted to an external public AI cloud. This makes the architecture potentially useful for government departments, banks, healthcare institutions, defence-related organisations and other sectors where data sovereignty, confidentiality and control over information are critical.
Ziroh has consequently developed industry-focused applications for sectors including healthcare, banking, finance, education and retail. Hospitals, for example, could potentially deploy local AI systems for clinical documentation and workflow support without sending patient information outside their own infrastructure. Banks could similarly use locally hosted models for document processing, customer-service applications, risk analysis and fraud-related workloads while retaining greater control over sensitive financial information.
The September 2026 launch represents the latest stage in the development of a technology that Ziroh first publicly unveiled at IIT Madras in April 2025. During that earlier demonstration, the company showed that foundation models could perform inference and certain other operations on CPUs. Ziroh had already optimised models from families including DeepSeek, Qwen, Llama and Phi while working with IIT Madras on benchmarking and research related to accessible artificial intelligence computing.
Since then, the platform has expanded substantially. The latest version supports multiple types of AI models and includes capabilities such as authentication, authorisation, telemetry, semantic memory and developer-friendly APIs. These improvements indicate that Ziroh is moving Kompact AI from an experimental computing concept toward a broader enterprise AI infrastructure platform.
The architecture could also support more distributed artificial intelligence systems. Instead of routing every request to a central AI data centre, separate Kompact AI deployments could operate at factories, hospitals, offices, banks and remote facilities. Processing information closer to where it is generated could reduce data movement, improve privacy and support applications that require faster local response times.
This distributed model also aligns with the growing international interest in sovereign AI, where countries and organisations seek greater control over the infrastructure, data and models used for artificial intelligence. India is simultaneously expanding access to high-performance AI computing, encouraging indigenous foundation models and seeking to build domestic capabilities across the AI technology stack. Technologies that allow advanced models to run on widely available hardware could therefore complement the country’s broader efforts to democratise AI computing.
Cost is likely to be one of the most closely watched aspects of Kompact AI. Ziroh has argued that selected workloads could achieve very large reductions in computing costs compared with GPU-heavy deployments, particularly where existing CPU infrastructure can be reused. The company has cited potential cost reductions of 80–90% for certain applications, although such figures will vary significantly depending on workload, model size, hardware and performance requirements.
The importance of Kompact AI therefore does not rest on eliminating GPUs from artificial intelligence. Instead, its significance lies in challenging the assumption that every AI workload needs specialised accelerator hardware. If CPU-native inference can provide commercially useful performance across a broad range of enterprise applications, organisations could make greater use of the enormous amount of general-purpose computing infrastructure already installed around the world.
For an Indian deep-tech company, the project is also notable because it addresses one of the foundational layers of artificial intelligence rather than simply building another application on top of existing global platforms. Ziroh Labs is attempting to rethink how AI models are executed and deployed, potentially making sophisticated artificial intelligence more accessible to organisations that cannot justify large investments in dedicated GPU infrastructure.
With the global rollout of Kompact AI, the company is effectively betting that the future of artificial intelligence will depend not only on building increasingly powerful specialised processors, but also on finding better ways to use the billions of CPUs that already exist across data centres, businesses and edge computing systems worldwide.
You may also like
-
Mysuru Deep-Tech Firm Vigyanlabs Launches Waterless FEMTO Sovereign AI Platform
-
Paras Defence Expands Indigenous Imaging Supply Chain With ₹8.63-Crore Precision Components Orders
-
Indian Navy Moves to Acquire Indigenous Arudhra 4D Surveillance Radars
-
Fresh ALH Procurement Approval Opens New Production Opportunity for HAL
-
NALCO Ties Up With Emirates Global Aluminium for Advanced 0.5-MTPA Angul Smelter Expansion