Grok 4.7 Is Now in Felo Search: Long-Horizon Answers, $2/$6
Grok 4.7 is live in Felo Search: 500K context, $2/$6 per million tokens, and long-horizon reasoning built for research that takes hours.
SpaceXAI shipped Grok 4.7 on September 21, and it is built for a specific kind of question: the one that takes an hour of digging rather than a single lookup. It is live in Felo Search now, alongside three other models that landed the same week.
The headline number sits on the output side. Grok 4.7 costs $2 per million input tokens and $6 per million output tokens, matching Grok 4.6's price while the vendor claims twice the speed of comparable models. A search answer in Felo is not one model call. It reads sources, checks them against each other, writes the synthesis, and goes back when they disagree. Output price is what a long investigation actually spends.

What Grok 4.7 Actually Is
SpaceXAI (the company formerly known as xAI) positions Grok 4.7 as its "most powerful model for coding and knowledge work." The interesting part of that description is the second half. Coding had been the focus of the previous three Grok releases; this one was trained with a longer reinforcement-learning run weighted toward tasks that take several hours to finish.
That training shows up in two places you can feel from a search box: self-verification and long-context management. The model is trained to check its own intermediate work rather than committing to the first plausible path, and to keep track of what it established twenty steps ago.
| Spec | Grok 4.7 |
|---|---|
| Model ID | grok-4.7 |
| Maker | SpaceXAI |
| Released | September 21, 2026 |
| Context window | 500,000 tokens |
| Input | Text, image |
| Output | Text |
| Reasoning effort | low / medium / high (default) / xhigh |
| Price | $2 per million input, $6 per million output |
| Cached input | $0.50 per million tokens |
| Knowledge cutoff | June 2026, with supplemental training through August 2026 |
The 500,000-token window is the smallest of the four models that arrived this week. It is also more than enough for a search task: a hundred long documents, a full quarter of earnings filings, or an entire research thread with every source still in context. Where the bigger windows matter is when you want the model to hold a repository or a year of documentation without a retrieval step in between. That is the case for the 1M-token models in the table below.
Why This Matters for AI Search
A search answer has two failure modes. The first is not finding the source. The second is finding the source, misreading it, and stating the misreading with confidence. The second one is worse, because it looks exactly like a correct answer.
Long-horizon training attacks the second failure mode. When a model is rewarded for finishing multi-hour tasks correctly, the cheap strategies stop working. Pattern-matching on the first plausible document, ignoring a source that contradicts the emerging answer, dropping a constraint established early: none of them survive a task that runs for hours. What survives that training is the habit of going back and checking.
Grok 4.7's published scores point the same direction. On GDPval, which measures real-world professional deliverable quality, the vendor reports an Elo of 1,695, up from 1,605 for Grok 4.6. On AA Briefcase v1.1, a knowledge-work benchmark, it moved from 1,546 to 1,657. HealthBench Professional went from 48.5% to 56.7%, and the Harvey Legal Agent Benchmark from 15.8% to 19.6%.
Those are vendor-reported figures, evaluated at different effort settings and harnesses, so treat them as a claim rather than a result. The comparison worth trusting is the one within the same family: 4.7 beats 4.6, consistently, on the work that resembles research.
How to Use Grok 4.7 in Felo Search
1. Open the model picker. Start a search in Felo and select Grok 4.7. Nothing else about your workflow changes: same box, same sources, same workspace.
2. Ask the question you would normally break into four. Long-horizon models earn their price on ambiguous, multi-constraint requests. Instead of "what is the EU AI Act timeline" try "which EU AI Act obligations land first for a company selling a general-purpose model into the EU, and which of them have implementing guidance published yet." One prompt, several sub-questions.
3. Let it disagree with itself. Follow up with "where do your sources conflict on this, and which one is more authoritative?" A model trained on self-verification handles the challenge better than one trained to produce a confident single answer.
4. Switch effort level when the question changes shape. Grok 4.7 exposes four reasoning-effort settings. Medium covers ordinary research; raise it when the question has a long chain of dependencies, and drop it when you are summarizing something you already trust.
5. Push the answer forward. The same context window holds the research and the deliverable. Ask for the memo, the deck outline, or the email. The source material is still in front of it.
The Other Three Models Live in Felo Search
Grok 4.7 did not arrive alone. Four frontier models shipped within about thirty-six hours of each other, and all four now run in Felo Search.
| Model | Released by | Price per million tokens | Context | Built for |
|---|---|---|---|---|
| Grok 4.7 | SpaceXAI | $2 in / $6 out | 500K | Long-horizon reasoning and knowledge work |
| Claude Opus 5.5 | Anthropic | $4 in / $20 out | 1M | Answers that survive verification |
| GPT-6 Luna | OpenAI | $0.10 in / $0.50 out | 1.05M | High-volume search at the lowest cost per call |
| GPT-6 Sol | OpenAI | $2 in / $10 out | 1.05M | Agent loops and repeat calls at mid-tier price |
The spread between them is roughly twenty times on input price and fifty times on the cost of a full investigation. Choosing per question is now the single biggest lever on both answer quality and how much research you can afford to run.
What to Try First
A question with a contested answer. Ask about something where serious sources disagree: a supplement's efficacy, a migration's real cost, a market-size estimate. Then ask Grok 4.7 to separate what the sources agree on from what only one of them claims. That is the shape of task the long-horizon training targets.
A document you have to argue with. Paste a contract, a policy, or a specification and ask what it does not cover. Absence-of-evidence questions are where a model that stops at the first plausible reading fails most quietly.
A cross-source reconciliation. Give it two filings from competing companies and ask for the comparison an analyst would make, with the number each one is quietly avoiding.
A forecast with constraints. Long-horizon models handle "here are eight constraints, find the plan that satisfies all of them" better than they handle open-ended brainstorming.
Try Grok 4.7 in Felo Search
Grok 4.7 is live in Felo Search today, with a 500,000-token context window, four reasoning-effort settings, and $2/$6 pricing. It costs the same as the model it replaces, which means the upgrade is free at the token level.
Open Grok 4.7 in Felo Search →
This post is also available in 简体中文, 日本語, 한국어, 繁體中文, हिन्दी, Français, العربية, Русский, اردو, Bahasa Indonesia, Deutsch, Tiếng Việt, Türkçe, Italiano, ไทย, Español, বাংলা and Português.