Open benchmark · August 2026

Tool calling,
measured in Malay.

A focused look at how four language models understand Malay prompts, function descriptions, and parameter descriptions across BFCL v3.

1,040test cases
4models compared
4Malay-focused categories
0tools executed

01 / Overall

The leaderboard

Exact local scoring over the same 1,040 cases. Higher is better.

Loading benchmark results…

What this measures

Tool choice, arguments, restraint, and conversational replies.

02 / Categories

Where models differ

Choose a category to compare the same models on a narrower skill.

Category

Simple tool calls

One function, one precise set of arguments.

03 / Exact results

Every score

Passed cases over scored cases. No hidden weighting.

Rank Model Overall Simple Multiple Irrelevance Chatable

04 / Model answers

Read every response

Eyeball the Malay, spot Indonesian phrasing, and compare all four models on the same prompt.

Download JSON ↓

Loading all model answers…

Answers are rendered from the original Markdown; open Raw Markdown to inspect the exact output.

Loading model answers…

Method note

A practical harness score,
not an official BFCL submission.

The focused suite covers simple, multiple, irrelevance, and chatable. It scores the first assistant response and never executes benchmark tools.

Results compare full serving paths. ILMU used a direct OpenAI-compatible endpoint; the other deployments used the Pixel Harness command adapter.

Read the full methodology