Sept 24, 2026 • Kefuro News
📦 Kefuro-30B enters training on real customer support outcomes

Read the plan

Introducing Kefuro Service beyond chat

The First Customer Service Model, Built For Every Business

Get Early Access
KF.AI.0S1
GPT
GPT
general chat
✣ chat
RAG
RAG
retrieval bots
✣ faq
KEFURO
KEFURO
trained on outcomes
✣ 30B✣ gate
Monthly cost at our volume
GPT-5.4 API$50,000
Kefuro · 2×RTX PRO 6000$2,100
Kefuro reply latency0.2–0.5s

≈15.4B input + 0.77B output tokens / month

Clock Tool 1.1
—
—
Kefuro
Version 0.01
© 2026 Kefuro
Made in San Jose. For sellers everywhere.

We Trained On Outcomes, Not Opinions

Not A Chatbot

General chat models are tuned to please whoever is typing. Support needs something else: the right answer, inside the company's policy, in the customer's language, every time.

A Specialist Model

Kefuro learns from millions of real support conversations and what happened next: resolved, escalated, refunded, or rated badly. Good outcomes teach it; bad ones become its negative examples.

Two Gates

Every reply passes a decision model before and after it is written. If the answer is not grounded in the account data or breaks policy, Kefuro escalates instead of guessing.

24x Cheaper, 0.3s Replies.

Cost: self-hosted estimate vs. GPT-5.4 list price at our volume. Latency: measured on one RTX PRO 6000, non-reasoning mode, internal preview.

How A Reply Gets Made

  1. 01

    Hard rules

    Safety, legal threats and "let me talk to a human" go straight to a person. No model decides those.

  2. 02

    Kefuro-Gate, before

    A calibrated decision model reads the ticket: intent, tool to call, whether to escalate, which tier should answer.

  3. 03

    Kefuro-30B writes

    The reply model answers from the customer's account data and the company's knowledge base, in the customer's language.

  4. 04

    Kefuro-Gate, after

    The same gate checks the draft: grounded, within policy, on-brand. Fail it and the reply is retried on a larger tier, then escalated.

The Kefuro Family

In training

Kefuro-30B

Reply model. 30B-parameter mixture-of-experts, about 3B active per token, built on NVIDIA Nemotron 3.5 Lightning and post-trained on support outcomes.

In training

Kefuro-Gate

Decision model. One forward pass returns calibrated probabilities for intent, escalation, tool choice and reply quality. Built on the open Laya model.

Coming soon

Kefuro-Bench

An open benchmark for customer service: multilingual tickets from many industries, policy traps and must-escalate cases, scored on outcomes rather than style.

Built for Retail & ecommerce · SaaS & subscriptions · Travel & hospitality · Financial services · Telecom & utilities · Local services

Our first training data comes from ecommerce support; other industries are added as partners join.

Languages English · Español · Français · Deutsch · Italiano · 日本語

Get Early Access

Early-access partners get the first Kefuro checkpoints, a seat in the evaluation, and a say in which industries we cover next.

early-access.form

Prefer email? support@flatkey.ai

Questions

What is Kefuro?

Kefuro is a family of models built only for customer service: a reply model, a decision model that checks every reply, and an open benchmark. It is meant for any business that answers customers, not one vertical.

Why not just use GPT or Claude?

They are excellent general models, but support is a narrow job with its own failure modes: invented dates and amounts, promises the policy does not allow, the wrong language. A specialist trained on real outcomes makes fewer of those mistakes and costs far less to run.

Which industries does it cover?

The goal is every business that runs customer service. We start from ecommerce, where our own support data comes from, and add industries such as SaaS, travel, financial services and telecom with partners who contribute data and evaluation.

Where does it run?

On GPUs we rent and operate ourselves. Customer conversations never pass through a third-party model API.

What does "24x cheaper" mean exactly?

At our current volume, the same traffic costs about $50,000 a month on GPT-5.4 list prices and about $2,100 a month on two self-hosted RTX PRO 6000 GPUs. It is an estimate from our own measurements, not a guarantee for other workloads.

Is it available now?

Not yet. Kefuro-30B and Kefuro-Gate are in training. Early-access partners will get the first checkpoints and a seat in the evaluation.

Will the weights be open?

Kefuro-Bench will be public. We plan to publish model weights on Hugging Face once they pass our own benchmark; the license will be announced with the release.