Two agents, one experiment: feel the difference
Two live agents are running in a sandbox right now, both reachable on WhatsApp. They have the same tools and the same Malaysian datasets. The only difference is the brain: one runs a frontier-class model, the other a budget model. Message both, ask the same question, and compare what comes back. This page is your guide while the demo runs.
The two agents
Agent GAJAH is the heavyweight. It runs a frontier-class model (OpenAI GPT-5.6 class), priced by OpenAI at $5.00 per million input tokens and $30.00 per million output tokens.
Agent KANCIL is the lightweight. It runs a budget model (DeepSeek v4 Flash), priced by DeepSeek at $0.14 per million input tokens and $0.28 per million output tokens.
| Agent | Model | Input, per million tokens | Output, per million tokens |
|---|---|---|---|
| GAJAH | Frontier-class (OpenAI GPT-5.6 class) | $5.00 | $30.00 |
| KANCIL | Budget (DeepSeek v4 Flash) | $0.14 | $0.28 |
Both agents carry identical tools and identical data. Only the brain differs. That is the whole point of the experiment: any difference you feel in the replies is the model, not the plumbing.
What to try
Ask anything about the Malaysian market. If you want a starting point, these questions exercise the datasets well:
- Compare median household income in Johor Bahru and Shah Alam.
- How many cafes are there in Penang?
- Which district in Selangor has the highest spending power?
- Profile the businesses within 2km of our venue.
- Where in Malaysia are pharmacies growing fastest?
The most useful move: send the same question to both agents, then compare. Look at three things. Depth: which answer goes further into the data? Caution: which agent flags what it does not know? Speed: which one comes back first? The differences are not subtle.
The rules of the sandbox
These rules are not decoration. They are a small, live example of the governance principles on the governance slide: constrain what an agent can reach, log what it does, and put a time limit on it.
- Replies take 30 to 60 seconds, and longer when the room is busy. The agents plan, query and check before answering; that takes time.
- Both agents are sandboxed on isolated infrastructure, with read-only access to public Malaysian datasets.
- They cannot act on the world. No email, no spending, no reach into any company system.
- Conversations are logged for the demonstration and deleted after. Do not send personal or confidential information.
- The agents stay live for one week after the workshop, so you can keep testing from your desk.
What you are feeling
Two gaps at once. The capability gap: the frontier model tends to reason further, hedge more honestly and handle messier questions. The cost gap: the budget model does respectable work at a fraction of the price. Neither agent is simply better. The lesson for an insights or marketing leader is that model choice is a management decision, not a technical one: match the model to the job, and pay for frontier reasoning only where the job demands it.
Sources
| Source | Supports | URL |
|---|---|---|
| OpenAI pricing | GAJAH model pricing: $5.00 input, $30.00 output, per million tokens | developers.openai.com/api/docs/pricing |
| DeepSeek pricing | KANCIL model pricing: $0.14 input, $0.28 output, per million tokens | api-docs.deepseek.com/quick_start/pricing |
BASIC · Agentic AI Workshop · aiagent.research.my