Flower Endeavor 1.0 arrives with a number that will get attention: 98.2 on HumanEval, higher than the four comparison models in Flower’s launch table. It also offers something many frontier APIs do not, a route to run the model inside infrastructure an organization controls.
The preview is not ready for a normal model-buying decision. The launch post and five-page datasheet do not state pricing, parameter count, context-window length, hardware requirements, or license terms. Strong benchmark scores can justify a pilot. They cannot fill those five procurement fields.
Flower introduced Endeavor 1.0 on September 1 as a limited preview for selected organizations and partners. Access is by request through a Flower-managed service or a private deployment.
The benchmark table is strong but not a sweep
Flower reports four completed core evaluations. Endeavor leads HumanEval. It ties GPT-5.6 Sol and Claude Fable 5 on AIME 2026. GPT-5.6 Sol is higher on GPQA and IFEval. These are company-reported results, and Flower itself warns that a small benchmark set is not enough to establish real-world quality.
| Benchmark | Endeavor 1.0 | GPT-5.6 Sol | Claude Fable 5 | Kimi K3 | Nemotron 3 Ultra |
|---|---|---|---|---|---|
| GPQA | 92.0 | 94.1 | 92.6 | 93.5 | 86.7 |
| HumanEval | 98.2 | 95.1 | 97.0 | 96.3 | 96.3 |
| IFEval | 94.1 | 95.9 | 91.7 | 92.8 | 91.9 |
| AIME 2026 | 99.9 | 99.9 | 99.9 | 96.7 | 94.2 |
The honest summary is narrower than “best frontier model.” Endeavor wins one listed benchmark, ties the top score on one, and trails GPT-5.6 Sol on two. It beats Nemotron 3 Ultra across all four and Kimi K3 across three of four. AIME’s three-way 99.9 also suggests the test may be saturated for this comparison.
Private deployment is the part competitors cannot copy with a price cut
Flower offers two routes. The managed service handles deployment, scaling, and model operations. Private deployment keeps the model and workload inside the customer’s environment, with Flower support. The company says applications can call Endeavor through familiar model APIs and response formats.
That choice can matter more than a one-point benchmark difference. A private deployment can reduce provider dependence, keep sensitive workloads inside an approved environment, and let a team hold a stable version while it validates upgrades. It can also move capacity planning, patching, observability, and incident ownership onto the buyer.
Our review of Tencent Hy4’s 214 GiB deployment reality shows why “available to run yourself” is not a deployment plan. The hardware footprint, serving stack, quantization, throughput target, and concurrency model determine whether control is affordable.
Five blank fields block a costed comparison
| Missing field | Why it changes the decision | Question for the preview |
|---|---|---|
| Pricing | Determines managed cost per task and private support economics | What are input, output, hosting, and minimum-commit charges? |
| Model size | Shapes memory, loading, replication, and upgrade work | What is the parameter count and active footprint? |
| Context window | Limits repository, document, and long-agent workloads | What is the supported context and its long-context pricing? |
| Hardware | Turns private deployment into an actual capacity plan | Which accelerators, memory, throughput, and concurrency are supported? |
| License and weights | Defines what “operate independently” legally and technically means | Are weights delivered, and what use, modification, and redistribution rights apply? |
The wording around open foundations deserves care. Flower says Endeavor builds on capabilities from leading open-weight models and adds its own training, integration, behaviors, and specialist work. It does not name the exact base models in the reviewed release. Do not infer a specific architecture, license, or downloadable checkpoint from the phrase “open-weight ecosystem.”
FlowerBench is promising evidence, but the launch gives no Endeavor task table
FlowerBench runs opt-in enterprise tasks inside participating organizations, keeping proprietary data and internal context in place while sharing sanitized results. Flower says those signals influenced Endeavor’s evaluation design, post-training, agent harnesses, and system improvements.
That is more relevant than another academic multiple-choice score. The launch does not publish an Endeavor result table across FlowerBench tasks. Buyers should ask for task success, verifier failures, time, tokens, cost, tool errors, recovery rate, and variance across repeated runs on work similar to their own.
A preview test should compare systems, not chat answers
- Choose five workflows with private context, tools, strict deliverables, and a measurable finish line.
- Run Endeavor through Flower-managed infrastructure and, if available, the proposed private stack.
- Use the same prompts, tool schemas, budgets, and stop conditions against the current production model.
- Score completion, verifier pass rate, human repair minutes, wall time, tokens, and total cost.
- Repeat each workflow enough times to expose variance and recovery failures.
The winning model is the one that finishes the work at an acceptable cost and failure rate. Our model selection guide explains why benchmark rank should be only one input to that assignment decision.
My verdict: request a pilot, not a migration
Flower Endeavor 1.0 has enough signal to justify a serious evaluation. The reported HumanEval result is strong, the overall table is competitive, and private deployment is a meaningful option for organizations that cannot build around one closed endpoint forever.
It does not yet have enough public information for a migration decision. Pricing, model size, context length, hardware requirements, and license terms define the real deployment. Ask for those fields, run private workflows with strict verifiers, and calculate the human repair cost. Frontier-class is a claim. A repeatable workload result is evidence.
Read the primary material
- Read Flower’s Endeavor 1.0 launch post.
- Review the current Endeavor model page and datasheet.
- Read the FlowerBench methodology and pilot discussion.
Checked September 1, 2026. Scores, access, deployment options, and product descriptions are reported by Flower. The missing-field audit and pilot design are Musthave.ai analysis. No independent Endeavor benchmark replication was available in the reviewed sources.