Top-line takeaway
Open benchmark · August 2026
Tool calling,
measured in Malay.
A focused look at how four language models understand Malay prompts, function descriptions, and parameter descriptions across BFCL v3.
01 / Overall
The leaderboard
Exact local scoring over the same 1,040 cases. Higher is better.
Loading benchmark results…
What this measures
Tool choice, arguments, restraint, and conversational replies.
02 / Categories
Where models differ
Choose a category to compare the same models on a narrower skill.
Category
Simple tool calls
One function, one precise set of arguments.
03 / Exact results
Every score
Passed cases over scored cases. No hidden weighting.
| Rank | Model | Overall | Simple | Multiple | Irrelevance | Chatable |
|---|
04 / Model answers
Read every response
Eyeball the Malay, spot Indonesian phrasing, and compare all four models on the same prompt.
Loading model answers…
Method note
A practical harness score,
not an official BFCL submission.
The focused suite covers simple, multiple, irrelevance, and chatable. It scores the first assistant response and never executes benchmark tools.
Results compare full serving paths. ILMU used a direct OpenAI-compatible endpoint; the other deployments used the Pixel Harness command adapter.
Read the full methodology