We started at 34 billion tokens. After moving to a detached system architecture, consumption dropped to around 1.5 billion tokens. With process segmentation and modular refactoring, we reduced that by roughly 50% from the 1.5 billion baseline, bringing effective token load below one billion.
That is not optimisation for show. That is measurable unit economics. Most AI cost leakage happens because the model is forced to reprocess the same system, the same process, the same logik, and the same workflow every time a small change is made.
But when the system is detached and each process is clearly segmented with its own dictionary, the AI reads only what it needs. Modify the payment flow? Read the payment process only. Modify order checking? Read the order process only. No full-system scanning. No token bonfire.
This is where Malaysian PKS can win. AI should not be reserved for companies with large servers and large budgets. AINNA's approach is about precise resource allocation, predictable OPEX, and stronger return on AI spend. The future is not just bigger models. The future is smarter systems and cleaner asset management.


