Table of Contents
GPT-5.6 Luna Terra Sol: Why AI Models Come in Sizes
GPT-5.6 Luna Terra Sol explained for parents: what each size costs, what each is good at, and why the right-size AI tool beats the biggest one for homework.
Nobody takes an 18-wheeler to pick up milk. That is the whole idea behind GPT-5.6 Luna, Terra, and Sol, the three sizes OpenAI released on July 9, 2026: same family, different engines, priced five times apart. Luna costs $1 per million input tokens, Terra $2.50, Sol $5. And on the independent Artificial Analysis Intelligence Index, the gap between them is 51, 55, and 59 points, which is a real difference but nowhere near five times.
For parents, that arithmetic is the lesson. Choosing the right size tool is an engineering skill, and it is teachable at almost any age.
Key Takeaways
- GPT-5.6 launched July 9, 2026 in three sizes: Luna (budget), Terra (middle), Sol (flagship), after a limited preview on June 26 during a dispute over government restrictions.
- API prices are $1/$6, $2.50/$15, and $5/$30 per million input/output tokens for Luna, Terra, and Sol.
- Artificial Analysis measured Intelligence Index scores of 51, 55, and 59 and per-task costs of $0.21, $0.55, and $1.04; a 5x price spread buys about 8 index points.
- Sol scored 80 on the Coding Agent Index, which OpenAI notes is 2.8 points above Claude Fable 5 while using less than half the output tokens.
- For homework, the cheap tier is usually the right tier; what matters is whether the kid asks the model to explain or to answer.
What GPT-5.6 Luna, Terra, and Sol actually are
A model tier is one AI system offered at several capability-and-cost points, usually by training or serving models of different sizes. OpenAI launched GPT-5.6 on July 9, 2026 with three: Sol is the flagship, which OpenAI called its best coding model yet; Terra sits in the middle; Luna is the fast, cheap option. All three arrived in ChatGPT, Codex, and the OpenAI API. A limited preview had gone out on June 26, held back over what reporting described as Trump administration concerns about the model’s cybersecurity capabilities. A separate cyber-focused variant, GPT-5.6-Cyber, followed on August 10.
The names are new, the structure is not. Google runs Flash and Pro lines. Anthropic sells Fable and Opus. Every lab has landed on the same conclusion: most requests are easy, a few are hard, and charging one price for both wastes money.
Why do the sizes differ in skill? Two mechanisms, both worth explaining to a kid. First, parameter count: bigger models store more patterns from training, which helps on unusual problems. Second, how long the model is allowed to think. These systems can spend extra computation before answering, and more thinking costs more tokens. That is why the same model at “max reasoning” scores higher and bills more than the same model at low effort. Our explainer on test-time compute walks through the second mechanism in detail.
What the independent numbers show
OpenAI’s marketing claim, from Sam Altman, is that Sol is “54% more token efficient” on coding tasks. The specific version OpenAI’s developer account published: on the Artificial Analysis Coding Agent Index, Sol at max reasoning set a new high of 80 “while using 54% fewer output tokens and completing tasks in 57% less time than the next-highest-scoring model.” That is a real, checkable claim about output tokens, not a vague quality boast, which makes it unusually good material for a media-literacy lesson. We use it as one in our guide to reading AI marketing numbers with your kid.
The independent Artificial Analysis evaluation, published the same day, is where the tier comparison gets useful. Intelligence Index: Sol 59, Terra 55, Luna 51. Cost to run the index: $1.04, $0.55, and $0.21 per task. Sol also used about 15,000 tokens per index task versus GPT-5.5’s 16,000, a modest efficiency gain rather than a revolution.
Here is the part that matters for a family: the jump from Luna to Sol is eight index points for five times the price. For a kid asking about the water cycle, those eight points are invisible. For a teen debugging a 400-line Python project, they might not be.
How to Teach Your Kid About AI Model Sizes
Ages 5–8: The vehicle game
Lay out toy vehicles: a bike, a car, a dump truck. Give your kid three jobs: take one letter to the neighbor, take the family to the park, move a pile of dirt. Let them match vehicle to job, then ask why the bike is wrong for the dirt and the dump truck is wrong for the letter. Now say the words: “AI helpers come in bike-size, car-size, and truck-size too.” That is the whole concept, and five-year-olds get it in about ninety seconds.
Ages 9–12: Same question, two tiers
If you have access to a model picker (ChatGPT, Claude, and Gemini all have one), ask the same question on the cheapest and the most capable model. Use something with a right answer that needs a few steps, like “a train leaves at 2:40 and arrives at 6:15, how long was the trip in minutes?” Have your kid grade both answers and time them. Most of the time they will find no difference, which is the point. Then try something genuinely hard and watch the gap appear.
Ages 13+: Build a routing rule
Have your teen write a one-page “routing policy” for their own AI use: which tier for flashcards, which for essay feedback, which for code, and what triggers an escalation to the expensive tier. Then have them check it against the actual price table below and calculate a monthly estimate. Real engineers build exactly this; it is called a model router, and a version of it runs inside every AI app your family uses.
The question to ask: “Sol scores 8 points higher than Luna and costs five times as much. Name a task where that trade is worth it, and one where it is not.”
The tier table: speed, cost, and what each size is for
| Tier | Price per 1M tokens (in / out) | Artificial Analysis Intelligence Index | Cost per index task | Best for |
|---|---|---|---|---|
| GPT-5.6 Luna | $1 / $6 | 51 | $0.21 | Flashcards, definitions, summaries, quick quizzing |
| GPT-5.6 Terra | $2.50 / $15 | 55 | $0.55 | Study help, essay feedback, multi-step word problems |
| GPT-5.6 Sol | $5 / $30 | 59 | $1.04 | Coding projects, hard proofs, long research tasks |
| GPT-5.6 Sol (max reasoning) | $5 / $30 plus reasoning tokens | 80 on Coding Agent Index | Higher | Agentic coding where correctness beats cost |
One caveat on the last row: “max reasoning” is a setting, not a separate product. The price per token stays the same; the bill rises because the model generates more thinking tokens, which are billed as output. A parent who sees a usage limit hit on Wednesday is usually looking at a kid who left reasoning on maximum.
What this means for learning, not just for bills
The tier question and the learning question are separate, and it is worth keeping them apart.
The OECD’s PISA 2025 results, released September 8, 2026 and covering roughly 760,000 15-year-olds, found that 46% of students in OECD countries use AI chatbots at least weekly to help them learn. Students who used a chatbot about weekly specifically for learning posted the highest science scores of any group, around 500 points, while students who used AI daily to draft written assignments scored 481 against 509 for those who never did, a 28-point gap after socioeconomic adjustment. OECD education director Andreas Schleicher’s framing was that technology can undermine learning when it “short-circuits the productive struggle of learning.”
Note what is not in that finding: nothing about model size. No study shows a more capable model producing more learning. The best-documented AI tutoring win, Kestin et al. (2025) in Scientific Reports, a randomized trial with 194 Harvard physics students, came from careful instructional design, expert-written scaffolds, and forced step-by-step reasoning, not from a bigger model.
So when your kid argues for the expensive tier, the honest answer is: for coding, maybe; for learning, the evidence is thin.
What to actually do at home
Make the cheap tier the default
Set the app to the fastest, cheapest model and let your kid escalate deliberately when something fails. This is the reverse of how most people use these tools, and it teaches judgment instead of habit.
Name the escalation trigger out loud
“Escalate when you have tried twice and can explain what is going wrong” is a good rule. It forces a diagnosis before a tool change, which is how debugging works in every technical field.
Check the reasoning setting monthly
Effort and reasoning controls reset with app updates. A quick look at the settings screen once a month prevents the Wednesday usage-limit surprise.
Let them see the price table
Kids who know Luna costs a fifth of Sol start making cost arguments themselves, sometimes annoyingly. That is a skill transfer worth the annoyance. Pair it with our tokens and per-million pricing explainer.
What not to do
Do not treat “which model” as the safety question. Tier choice is a cost-and-capability decision. Safety lives in account settings, age policies, and the rules about when AI may be used at all, which is a separate conversation covered in the parents’ guide to summer 2026 models.
What to Watch For Over the Next 3 Months
- Week 4: Open the model picker with your kid and see which tier the app defaults to. Many apps auto-route and do not tell you.
- Month 2 red flags: Usage limits hit early in the week; a teen who insists only the flagship “works”; a sharp drop in how often they attempt a problem before asking.
- Month 3 self-check: Ask your kid to justify one escalation to the expensive tier from the past month. If they can name the specific failure that triggered it, the routing habit has stuck.
Frequently Asked Questions
What is the difference between GPT-5.6 Luna, Terra, and Sol?
They are three capability-and-price tiers of the same model family released July 9, 2026. Luna is cheapest and fastest ($1/$6 per million tokens), Terra is the middle option ($2.50/$15), and Sol is the flagship ($5/$30), which OpenAI positions as its strongest coding model.
Is Sol worth five times Luna’s price for homework?
Usually not. Independent testing put Sol at 59 on the Artificial Analysis Intelligence Index versus Luna’s 51. For definitions, summaries, and quizzing, that difference rarely shows. For agentic coding, where Sol scored 80 on the Coding Agent Index, it can matter.
What does “54% more token efficient” actually mean?
OpenAI’s specific claim is that Sol at max reasoning reached a new high of 80 on the Artificial Analysis Coding Agent Index while using 54% fewer output tokens and taking 57% less time than the next-highest-scoring model. It is a claim about tokens and time on one benchmark, not a general statement about quality.
Which tier does ChatGPT use for my kid?
The app decides, often automatically, and it changed several times over 2026. Open the model picker together to see. Accounts placed in ChatGPT for Teens, launched August 18, 2026, also apply separate behavioral rules regardless of tier.
Why do AI companies make several sizes at all?
Because serving a large model is expensive and most requests are easy. Tiering lets a company charge less for the easy majority. The same logic gives you economy and express shipping, or a bike and a truck in the same garage.
About the author
Ricky Flores is the founder of HiWave Makers and an electrical engineer with 15+ years of experience building consumer technology at Apple, Samsung, and Texas Instruments. He writes about how kids learn to build, think, and create in a tech-saturated world. Read more at hiwavemakers.com.
Sources
- TechCrunch. (2026, July 9). “OpenAI launches its new family of models with GPT-5.6.” https://techcrunch.com/2026/07/09/openai-launches-its-new-family-of-models-with-gpt-5-6/
- Artificial Analysis. (2026, July 9). “GPT-5.6 benchmarks across Intelligence, Speed and Cost.” https://artificialanalysis.ai/articles/gpt-5-6-has-landed
- OpenAI. (2026, July 9). “GPT-5.6: Frontier intelligence that scales with your ambition.” https://openai.com/index/gpt-5-6/
- Wikipedia. (2026). “GPT-5.6.” https://en.wikipedia.org/wiki/GPT-5.6
- Kestin, G., Miller, K., Klales, A., et al. (2025). “AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting.” Scientific Reports. https://www.nature.com/articles/s41598-025-97652-6
- OECD. (2026, September 8). “PISA 2025: Students’ reading and mathematics performance declined sharply across the OECD.” https://www.oecd.org/en/about/news/press-releases/2026/09/pisa-2025-students-reading-and-mathematics-performance-declined-sharply-across-the-oecd.html
- OpenAI. (2026, August 18). “ChatGPT for Teens.” https://openai.com/index/chatgpt-for-teens/