Kefuro-30B
Reply model. 30B-parameter mixture-of-experts, about 3B active per token, built on NVIDIA Nemotron 3.5 Lightning and post-trained on support outcomes.
Introducing Kefuro Service beyond chat
≈15.4B input + 0.77B output tokens / month
Not A Chatbot
General chat models are tuned to please whoever is typing. Support needs something else: the right answer, inside the company's policy, in the customer's language, every time.
A Specialist Model
Kefuro learns from millions of real support conversations and what happened next: resolved, escalated, refunded, or rated badly. Good outcomes teach it; bad ones become its negative examples.
Two Gates
Every reply passes a decision model before and after it is written. If the answer is not grounded in the account data or breaks policy, Kefuro escalates instead of guessing.
Cost: self-hosted estimate vs. GPT-5.4 list price at our volume. Latency: measured on one RTX PRO 6000, non-reasoning mode, internal preview.
Safety, legal threats and "let me talk to a human" go straight to a person. No model decides those.
A calibrated decision model reads the ticket: intent, tool to call, whether to escalate, which tier should answer.
The reply model answers from the customer's account data and the company's knowledge base, in the customer's language.
The same gate checks the draft: grounded, within policy, on-brand. Fail it and the reply is retried on a larger tier, then escalated.
Reply model. 30B-parameter mixture-of-experts, about 3B active per token, built on NVIDIA Nemotron 3.5 Lightning and post-trained on support outcomes.
Decision model. One forward pass returns calibrated probabilities for intent, escalation, tool choice and reply quality. Built on the open Laya model.
An open benchmark for customer service: multilingual tickets from many industries, policy traps and must-escalate cases, scored on outcomes rather than style.
Built for Retail & ecommerce · SaaS & subscriptions · Travel & hospitality · Financial services · Telecom & utilities · Local services
Our first training data comes from ecommerce support; other industries are added as partners join.
Languages English · Español · Français · Deutsch · Italiano · 日本語
Early-access partners get the first Kefuro checkpoints, a seat in the evaluation, and a say in which industries we cover next.
Prefer email? support@flatkey.ai
Kefuro is a family of models built only for customer service: a reply model, a decision model that checks every reply, and an open benchmark. It is meant for any business that answers customers, not one vertical.
They are excellent general models, but support is a narrow job with its own failure modes: invented dates and amounts, promises the policy does not allow, the wrong language. A specialist trained on real outcomes makes fewer of those mistakes and costs far less to run.
The goal is every business that runs customer service. We start from ecommerce, where our own support data comes from, and add industries such as SaaS, travel, financial services and telecom with partners who contribute data and evaluation.
On GPUs we rent and operate ourselves. Customer conversations never pass through a third-party model API.
At our current volume, the same traffic costs about $50,000 a month on GPT-5.4 list prices and about $2,100 a month on two self-hosted RTX PRO 6000 GPUs. It is an estimate from our own measurements, not a guarantee for other workloads.
Not yet. Kefuro-30B and Kefuro-Gate are in training. Early-access partners will get the first checkpoints and a seat in the evaluation.
Kefuro-Bench will be public. We plan to publish model weights on Hugging Face once they pass our own benchmark; the license will be announced with the release.