
The end of expensive AI? Price shock at OpenAI: Why GPT-6 Sol and Luna are turning the AI market upside down – Creative image on the topic, with AI: Xpert.Digital
More performance, half the price: How OpenAI's new GPT-6 duo is putting pressure on the competition
Attack on Anthropic: OpenAI destroys the existing pricing logic with GPT-6 Sol & Luna
Intelligence at a bargain price: How GPT-6 is changing the economics of AI
With the surprise launch of its new GPT-6 Sol and GPT-6 Luna models, OpenAI is ushering in a new era in the artificial intelligence competition. Instead of simply showcasing new top-tier performance, the company is attacking its competitors – particularly Anthropic – where it matters most to the masses: price. By drastically halving usage costs compared to the previous generation, AI is finally transforming from an expensive luxury item into ubiquitous infrastructure. However, these new, extremely affordable pricing models also present a strategic pitfall: While automation suddenly costs a fraction of its former price, opening up enormous opportunities for business processes, it also threatens to drive a rapid increase in overall consumption due to new use cases. Those who want to emerge as winners in this new AI economy cannot blindly rely on the lowest token price, but must know how to cleverly orchestrate models in the future.
OpenAI's price attack: GPT-6 Sol and Luna shift the economics of AI
Cheaper intelligence will become a problem for everyone who still profits from scarcity
OpenAI is launching two new models, GPT-6 Sol and GPT-6 Luna, less than three months after the general release of the GPT-5.6 family, which are intended to mark the next stage in the mass market for artificial intelligence. The central message is at least as significant economically as it is technologically: increased performance should not come at a higher price, but rather with significantly lower usage costs. GPT-6 Sol costs two US dollars per million input tokens and ten US dollars per million output tokens via the API. For GPT-6 Luna, the prices are ten cents for input and 50 cents for output. Compared to the most recent promotional prices of the corresponding GPT-5.6 models, OpenAI is thus roughly halving the prices, with the reduction for Luna's output tokens being even slightly more than 58 percent.
This price movement is more than just a typical product update. It demonstrates that competition among major AI providers is increasingly shifting from simply chasing top-tier performance to a second level: The decisive factor is how much economically viable performance a model generates for a given amount of money. This is precisely where Sol and Luna come in. Sol is designed to handle demanding knowledge work, software development, and agent-based processes at costs previously associated with mid-range models. Luna, on the other hand, aims to perform routine tasks so cost-effectively that AI can be deployed in business processes where automation has previously been hardly worthwhile.
The provocative interpretation, therefore, is that OpenAI isn't simply selling better models at lower prices, but rather attempting to disrupt the existing pricing logic of the market. If high model quality no longer necessarily entails high variable costs, premium prices lose some of their justification. At the same time, the pressure on competitors increases to either offer their models at lower prices or to demonstrate their added value much more clearly. For companies, this development opens up new application possibilities. However, it also carries the risk that falling unit costs will lead to an uncontrolled expansion of usage, ultimately resulting in higher overall expenditures.
Two models for different value creation
The names Sol and Luna reflect a functional division of the model portfolio. GPT-6 Sol is the higher-performing mainstream model positioned below the top-of-the-line GPT-6 Astra. It is designed for tasks where accuracy, multi-level reasoning, reliable tool usage, and the ability to handle longer workflows are economically important. These include complex analyses, demanding programming tasks, automated research, the control of digital applications, and the support of professional knowledge work. GPT-6 Luna, on the other hand, is positioned as a particularly cost-effective and fast model for high-volume tasks. It is especially suitable for classification, extraction, summarization, simple coding tasks, preliminary checks, and standardized communication.
This tiered approach reflects an increasingly important realization: companies don't need the most powerful available model for every request. A large portion of business AI usage consists of repeatable tasks with limited complexity. In these cases, a very inexpensive model can deliver the greatest economic benefit. Only difficult, ambiguous, or high-risk cases need to be passed on to a more powerful model. Sol and Luna thus support an architecture in which requests are dynamically distributed based on difficulty, risk, and value.
For OpenAI, this segmentation makes strategic sense. The expensive flagship model Astra can continue to serve as a technological benchmark without its high cost hindering widespread adoption. Sol occupies the market for high-performance production systems, while Luna covers the price-sensitive mass market. In this way, OpenAI can simultaneously defend high margins in the premium segment, challenge competitors in the mid-range, and build enormous usage volumes in the low-price segment.
From the customer's perspective, however, this creates a new procurement challenge. The choice of a model should not be based solely on the list price. Crucial factors include which quality level is sufficient for a specific process, how often the model makes errors, how expensive it is to detect them, and the effort required for human oversight. A seemingly inexpensive model can become costly if its results frequently require rework. Conversely, a more robust model can be more economical if it reliably achieves the desired result with fewer iterations and less oversight.
Halving the prices changes the calculation
The price drop is particularly easy to understand with Sol. GPT-5.6 Sol recently cost four US dollars per million input tokens and 20 US dollars per million output tokens as part of a limited-time promotion. GPT-6 Sol halves both values to two and ten US dollars, respectively. With Luna, the input price drops from 20 to ten cents. The output price falls from 1.20 US dollars to 50 cents. The blanket statement that both models are exactly 50 percent cheaper therefore oversimplifies the actual structure: Luna's output price drops more significantly, which is particularly relevant for text-intensive applications.
Output tokens represent the larger cost factor in many applications. A system might receive a relatively short instruction but generate lengthy reports, code snippets, or structured datasets. It is precisely in such cases that the significant reduction in Luna's output price has a noticeable impact. With one million input tokens and one million output tokens, Luna costs a total of 60 cents. Sol, under the same simplified assumptions, costs twelve US dollars. This represents a factor of twenty difference between the two models. This gap clearly illustrates why intelligent model selection is becoming more important than simply committing to a single model.
For a company with 500 million input and 100 million output tokens per month, Sol would cost around $2,000 per month at standard pricing. Luna would cost approximately $100. This model calculation does not include tools, search access, data storage, vector databases, security audits, or development costs. Nevertheless, it illustrates the scale: at high volumes, the right model assignment can make a significant difference. The effect is even greater when recurring contextual information is read from the cache, thus incurring only a fraction of the normal input price.
The new prices simultaneously establish a new benchmark for the market. Claude Opus 5 was offered at $5 for input and $25 for output, while Claude Fable 5.1 was priced at $10 and $50, respectively. Sol significantly undercuts these list prices, and Luna is in a completely different price bracket. A simple token comparison isn't sufficient, as models require varying numbers of tokens and processing steps. Nevertheless, it will become more difficult for providers to charge several times the OpenAI price if they can't demonstrate a correspondingly greater benefit in real-world applications.
Accuracy is more valuable than a low token price
OpenAI is emphasizing improved factual accuracy in its announcement. According to an internal audit, GPT-6 Sol made approximately half as many factual errors as its predecessor. While this statement is economically significant, it shouldn't be reduced to the claim that the overall error rate has dropped from 20 to 10 percent. OpenAI is not publishing such absolute figures for this comparison. The audit was conducted using anonymized real-world conversations in which users had previously reported a factual error. These are therefore deliberately error-prone cases and not a representative sample of the entire user base.
The phrase "half as many errors" refers to a relative improvement within this specific test. It does not mean that Sol is now twice as reliable in every domain. Error rates depend heavily on the subject matter, its recency, data access, language, task definition, and permitted tools. A model can be very reliable with widespread general knowledge and simultaneously fail with rare industry information, current events, or precise figures. Companies should therefore understand the manufacturer's statement as an indication of progress, not a guarantee.
Nevertheless, even a relative halving of errors possesses considerable economic potential. In many production processes, the greatest costs arise not from model queries, but from monitoring, correction, and liability risks. If a system delivers a problematic result only in one out of every ten responses instead of every fifth, the testing effort does not automatically decrease by half. Full inspection may still be necessary. However, in less critical processes, the improved reliability can enable the transition from purely assistive use to largely automated processing.
Particularly noteworthy is OpenAI's claim that, under high computational intensity, GPT-6 Luna can achieve the level of GPT-5.6 Sol in this factual accuracy check at approximately one-hundredth of the cost. This statement refers to the factual accuracy tested and should not be extrapolated to overall model performance. Luna will not automatically perform as well as its predecessor Sol in every analysis, every code project, or every agent-based workflow. Nevertheless, the comparison demonstrates how quickly capabilities from expensive, top-tier models migrate to more affordable ones.
Professional work is becoming a price-performance competition
In the AutomationBench test, GPT-6 Sol, running at very high computational intensity, achieves a score of 33.2 percent. Claude Opus 5, at maximum settings, reaches 26.9 percent. The difference is 6.3 percentage points, not 6.3 percent. Relatively speaking, Sol is about 23 percent above Claude's score. At the same time, OpenAI puts the cost of Sol at 27 cents per task, while Opus 5 cost 11.1 times as much in this test. Sol therefore achieves a higher test score at only about nine percent of the comparison cost.
The benchmark maps complete workflows using 47 tools across sales, marketing, operations, customer service, finance, and human resources. This is more meaningful for business applications than purely knowledge-based questions because a productive agent must not only generate text but also operate systems, process intermediate results, and correctly connect multiple steps. At the same time, the absolute score of 33.2 percent remains sobering. Even the best result means that a large proportion of tasks are not completed fully. Progress is real, but a universally reliable digital workforce is still a long way off.
GPT-6 Luna improves by 5.4 percentage points compared to its predecessor at high computational intensity, while the cost per task is said to decrease by 58 percent. Here, too, linguistic precision is important. Percentage points describe the absolute difference between two performance values. An increase of 5.4 percent, on the other hand, would be a relative change. This difference can be significant when evaluating a model, especially if initial values are low.
Another test, Agents' Last Exam, evaluates long-term and economically relevant tasks from 55 sub-industry sectors. GPT-6 Sol achieves 56.4 percent at maximum computational intensity, surpassing, according to OpenAI, the best score of Claude Opus 5 while using 60 percent less cost per task. This underscores the direction but also the limits of development. A success score just above half is valuable for supervised assistance but often insufficient for unsupervised core processes.
Programming remains the toughest competitive field
In software development, the differences between the top-performing models are smaller. In the DeepSWE test, GPT-6 Sol achieves 68.8 percent efficiency at maximum computational intensity. Claude Fable 5 reaches 69.9 percent at very high settings. Sol thus lags behind by 1.1 percentage points, but is said to complete the task at approximately 80 percent lower cost. GPT-6 Luna achieves 66.6 percent, placing it close to Claude Opus 5 and Claude Fable 5 at medium computational intensity. According to OpenAI, Luna costs 93 percent less per task than Opus 5 and 96 percent less than Fable 5.
For companies, this scenario is more economically attractive than a narrow test victory. A model that delivers nearly the same success rate at a fraction of the cost can finance more attempts, additional tests, and automated cross-checks. Instead of paying for one expensive iteration, a system can generate several cost-effective solutions, run tests, and only select the best candidate. The lower unit costs can thus indirectly improve the quality of the entire development process.
However, this effect should not be overestimated. Repeated attempts are only useful if errors can be detected. With software, this is partially possible because compilation, automated tests, and static analysis provide objective feedback. Evaluation is more difficult with architectural decisions, security issues, or incomplete requirements. More generated code can then also mean more technical debt and additional testing effort.
OpenAI also refers to FrontierCode, a test that evaluates not only functional correctness but also the actual integrability of changes. This includes test quality, limited scope of changes, programming style, and adherence to project-specific rules. These criteria are crucial in practice. A code proposal can be formally correct yet still be unsuitable because it unnecessarily modifies large parts of a system or disregards operational standards. The economic value of a programming model, therefore, lies not in the amount of code generated, but in the number of reliably integrable changes per hour of work invested.
Computer control is becoming cheaper, but not risk-free
When operating graphical applications, GPT-6 Sol achieves 60.5 percent efficiency in the offline portion of OSWorld 2.0 with very high computational intensity. Claude Opus 5 achieves 60.3 percent efficiency at medium settings. OpenAI estimates Sol's cost per task at approximately 80 percent less. Luna, at maximum settings, is expected to surpass the average efficiency of GPT-5.6 Sol while costing only one-tenth of the price.
This development is significant for office automation, customer service, and back-office processes. Many companies still use older applications without modern programming interfaces. A model that understands screen content, operates buttons, and transfers information between systems can bridge these gaps. This creates a form of software-based automation that is more flexible than traditional rule-based robots.
However, even a value around 60 percent shows that unattended computer control remains risky. In lengthy workflows, a single wrong click is enough to cause the entire process to fail. Payments, access rights, data deletion, and changes to production systems are particularly critical. Therefore, tiered permissions, limited scopes of action, logging, and human approval at high-risk points are economically sound.
The real progress lies less in the fact that Sol has already completely replaced an employee. More importantly, the costs of experimental automation are decreasing. Companies can test more processes without having to implement every idea through an expensive integration project. Standardized agents can emerge from successful pilot projects; unsuitable cases can be discarded early on. The price reduction thus shortens the learning curve for user companies.
A new dimension of digital transformation with 'Managed AI' (Artificial Intelligence) - Platform & B2B solution | Xpert Consulting
A new dimension of digital transformation with 'Managed AI' (Artificial Intelligence) – Platform & B2B solution | Xpert Consulting - Image: Xpert.Digital
Here you will learn how your company can implement customized AI solutions quickly, securely and without high entry barriers.
A managed AI platform is your all-inclusive, worry-free solution for artificial intelligence. Instead of dealing with complex technology, expensive infrastructure, and lengthy development processes, you receive a ready-made solution tailored to your needs from a specialized partner – often within just a few days.
The key advantages at a glance:
⚡ Rapid implementation: From idea to ready-to-use application in days, not months. We deliver practical solutions that create immediate added value.
🔒 Maximum data security: Your sensitive data stays with you. We guarantee secure and compliant processing without sharing data with third parties.
💸 No financial risk: You only pay for results. High upfront investments in hardware, software, or personnel are completely eliminated.
🎯 Focus on your core business: Concentrate on what you do best. We take care of the entire technical implementation, operation, and maintenance of your AI solution.
📈 Future-proof & scalable: Your AI grows with you. We ensure continuous optimization and scalability, and flexibly adapt the models to new requirements.
More information here:
The impact of GPT-6 Sol and Luna on businesses
Intermediate storage is becoming an underestimated cost driver
In addition to the visible token prices, OpenAI improves the caching of recurring input. With so-called prompt caching, identical context fragments don't need to be completely reprocessed with every request. GPT-6 grants a 90 percent discount on previously read, cached input tokens. This is particularly relevant for agents, long conversations, and applications with extensive system instructions.
A business agent often receives the same information at every step: role description, security rules, data structures, tool definitions, manuals, or parts of the previous conversation. Without effective caching, this context is processed repeatedly at full cost. In multi-stage processes, the repeated context can represent a larger cost component than the actual new request. Therefore, a high cache hit rate impacts cost-effectiveness more than the list price would suggest.
OpenAI now allows developers to modify computational intensity and available tools without losing the previous context for caching. Furthermore, developers can set explicit breakpoints to specify which part of an input is reused. GitHub reports that these improvements reduced the proportion of prompt tokens that had to be freshly processed across billions of Copilot requests by more than 50 percent within a few months.
For practical purposes, this leads to a clear priority: cost optimization doesn't begin with the choice of model. It also encompasses the structure of inputs, the order of stable and variable content, the length of the context, and the number of unnecessary repetitions. A poorly designed application can be expensive despite using inexpensive models. Conversely, a well-planned architecture can strategically employ high-quality models without exceeding the budget.
Cheaper tokens do not automatically mean smaller bills
The AI inference market has been experiencing an extraordinary price decline for years. Between November 2022 and October 2024 alone, the cost to achieve roughly the same performance level as GPT-3.5 in a widely used natural language understanding test fell from $20 to seven cents per million tokens. This represented a decrease of more than 280 times. GPT-6 Sol and Luna continue this trend, albeit at a higher capability level.
Economically, a mechanism similar to the Jevons paradox is at play here. If the use of a resource becomes more efficient and cheaper, overall consumption does not necessarily decrease. Often, it increases because new applications become economically viable. A company that previously only used AI to generate a few high-quality texts can now analyze every customer request, classify every document, and automatically check many software changes at significantly lower prices. In this case, consumption increases faster than the price per unit falls.
Agent-based applications amplify this effect. A single user query can trigger ten or twenty model calls. The agent plans, searches for information, uses tools, checks intermediate results, and restarts if errors occur. Therefore, the visible query is not a suitable unit of cost. What matters are the total costs per successfully completed operation.
Companies should therefore not ask which model sells the cheapest tokens. They should measure how much a correct, timely, and verifiable result costs. This calculation includes model calls, search and tooling costs, infrastructure, human oversight, error correction, and potential damages. Only this comprehensive analysis reveals whether Sol or Luna is more economical.
OpenAI's attack on Anthropic is targeted
The announcement frequently compares Sol and Luna to Anthropic models. This is no coincidence. Claude is considered a strong provider for programming, long contexts, and demanding knowledge work in many companies and development teams. OpenAI is therefore not only targeting abstract peak performance, but also those application areas in which Anthropic has established a credible market position.
The price difference is clear. GPT-6 Sol costs two US dollars to buy and ten US dollars to issue per million tokens. Claude Opus 5 costs five and 25 US dollars respectively, and Claude Fable 5.1 costs ten and 50 US dollars respectively. On paper, Sol thus costs 60 percent less than Opus 5 and 80 percent less than Fable 5.1. GPT-6 Luna undercuts these prices even further.
The crucial part of the message, however, is not "cheaper per token," but "cheaper per task." OpenAI is trying to demonstrate that its models remain more cost-effective even when different computational intensity, response length, and tool usage are factored in. This is precisely why the company publishes costs per task on AutomationBench, DeepSWE, and OSWorld. This metric is fundamentally more meaningful than the list price alone.
Nevertheless, the comparisons are not entirely neutral. OpenAI conducts the analyses in its own research environment or via its own API. Some competitor data is taken from publicly available reports, and data for Claude Fable 5 is used when data for Fable 5.1 is unavailable. Furthermore, the cost analysis for AutomationBench with Fable 5.1 is incomplete because reliance on Opus 5 was not included in approximately 40 percent of the tasks. This makes the comparison more favorable for Anthropic, but also demonstrates how difficult it is to achieve a completely standardized measurement.
Manufacturer benchmarks provide evidence, but not judgments
Benchmarks are necessary to measure progress. However, they do not provide an objective overall score. Results depend on the test version, system settings, tools, computing budget, number of trials, and evaluation methods. Even a different computational intensity setting can significantly alter the value and costs. When comparing models with different budgets, an apparent difference in quality may simply be the result of additional computational effort.
With GPT-6, it's also important to note that many key statements originate from OpenAI's own publication. The company possesses detailed knowledge of its models but also has a sales interest. Therefore, the reported results should be taken seriously but supplemented by independent testing. Particularly small differences, such as the 1.1 percentage points between Sol and Claude Fable 5 in DeepSWE, may fall within a practical margin of error.
Even a high benchmark score doesn't guarantee good performance in a specific organization. Internal documents, specialized terminology, outdated software, specific compliance rules, and inconsistent data can all negatively impact results. Conversely, a model might perform better in a narrowly defined, well-documented process than in a general test.
The most robust procurement method is therefore a dedicated evaluation set based on real-world tasks. It should include common standard cases, challenging edge cases, and deliberately risky situations. Measurements should encompass not only accuracy and cost, but also processing time, stability, traceability, and necessary human intervention. A model change is only economically justified if this evaluation demonstrates a substantial advantage.
Clearer language reduces indirect process costs
OpenAI promises not only technical improvements but also a clearer communication style. Sol and Luna are intended to use less jargon, fewer unusual formulations, and fewer irrelevant details. Answers should be somewhat shorter without losing substance. This change may initially sound cosmetic, but it could have significant economic implications.
Unclear answers lead to follow-up questions, misinterpretations, and additional processing effort. In software development, a vague description can obscure what was actually tested. In customer communication, unnecessary technical jargon can reduce clarity. In analyses, long but content-free passages can increase the testing effort. Precise and concise communication therefore not only improves the user experience but also reduces process costs.
At the same time, style remains subjective. Shorter answers are not automatically better. For legal, technical, or security-critical topics, a detailed explanation may be necessary. Companies should therefore define their own guidelines for structure, tone, level of detail, and uncertainty labeling. The quality of communication arises from the interplay of the model, system instructions, and quality control.
Another improvement concerns honesty about actions performed. OpenAI reports lower rates of misleading statements about completed programming tasks. This is crucial for agents. A system that claims to have run tests or checked files when it hasn't creates a dangerous sense of security. Improvements in this area are more important than mere stylistic polishing.
Safety and liability remain outside the price tag
OpenAI explains that Sol and Luna have improved in alignment and security compared to their GPT 5.6 predecessors. This includes checks for misleading information in programming tasks, warning bypassing, and unauthorized interactions. These tests are intentionally difficult and do not represent typical usage. Rather, they demonstrate how models behave under problematic conditions.
For companies, a higher security score does not mean that technical and organizational safeguards can be dispensed with. Models require clearly defined rights, separate development and production environments, logging, access controls, and approvals for sensitive actions. When dealing with personal data, data protection, retention policies, and regional processing must also be considered. In regulated sectors, results must be traceable and responsibilities clearly defined.
Paradoxically, the low prices could create new risks. If Luna becomes so cheap that millions of automated decisions are possible, the reach of a systematic error increases. A faulty classifier, a flawed rule, or a distorted data set can then affect a very large number of cases. The cost per query decreases, but the potential total damage grows with the volume.
Companies should therefore not cut security and quality budgets proportionally to token prices. On the contrary: the more a model scales, the more important monitoring, sampling, and shutdown mechanisms become. Economic scaling without governance is not efficient, but rather a shifting of costs into the future.
The real innovation lies in the model orchestration
Sol and Luna suggest a multi-tiered architecture. Luna can handle large volumes of simple tasks while simultaneously assessing whether a case is unsafe or unusual. Sol handles more complex exceptions. Astra remains reserved for particularly difficult or business-critical tasks. Additionally, rules can define when human intervention is required.
Such a cascade combines low average costs with high peak performance. However, its success depends on the routing logic. If Luna fails to identify difficult cases, quality issues arise. If too many cases are preemptively routed to Sol or Astra, the cost advantage disappears. The classification of task difficulty thus becomes a core AI function in itself.
Different vendors can also be combined. A company doesn't have to commit entirely to OpenAI or Anthropic. It can distribute models based on task, region, availability, and price. Such a multi-vendor approach reduces dependencies and increases reliability. However, it also creates additional integration and testing effort because the models have different interfaces, tool formats, and behaviors.
In the long run, competition is unlikely to be decided solely between individual models. Platforms that automatically route requests, monitor costs, measure quality, and exchange models as needed will be crucial. Sol and Luna increase the benefits of such systems because the price differences between the tiers are large enough to make intelligent distribution economically attractive.
Winners and losers of the new prize round
Among the immediate winners are developers of AI applications with high usage volumes. Services for document processing, support, software development, data cleansing, and research can significantly reduce their variable costs. Startups benefit because they can conduct more experiments with less capital. Medium-sized companies gain access to capabilities that were previously only viable with larger budgets.
Users of specialized software could also benefit. If providers pass on lower model costs, AI features can be integrated into existing subscriptions without significantly increasing prices. However, it's more likely that some of the savings will remain with the software providers as margin. Competition will determine how much of this actually reaches the customer.
Providers whose business model relies primarily on reselling expensive model access are coming under pressure. As the underlying technology becomes cheaper and more powerful, the value of simple user interfaces and interchangeable add-on functions diminishes. Sustainable differentiation then arises from proprietary data, deep process integration, industry knowledge, security, and demonstrable results.
Anthropic and other model providers also face a strategic decision. They can lower prices, expand their core capabilities, or highlight particular advantages in security, usability, and enterprise deployment. Pure price competition favors providers with large infrastructure, high utilization, and access to capital. Smaller labs could focus more on open models, specialized niches, or sovereign deployment.
What companies should specifically check now
The first task is to break down existing AI costs by process rather than by model. Companies need transparency into how many input, output, and buffer tokens a complete process consumes. Additionally, tool calls, repetitions, and human review time must be tracked. Without this data, the economic benefits of switching cannot be reliably determined.
Real-world comparative tests between the previously used model, GPT-6 Sol, and GPT-6 Luna should follow. Luna is the natural candidate for high volumes and clearly structured tasks. Sol is suitable for complex analyses, longer processes, and demanding software applications. For particularly difficult tasks, a top-tier model can still be economical if it avoids repetition and errors.
A hasty, complete migration would be unwise. New models, despite higher average performance, may have different error profiles. Outputs may change in format, length, or tone, thereby disrupting downstream systems. A phased implementation with parallel execution, automated quality checks, and clear fallback options is advisable.
Contracts and technical architecture should also facilitate model switching. Vendor-specific features may be convenient in the short term, but increase switching costs later on. Standardized data formats, separate evaluation logic, and centralized logging help to test new models more quickly. In a market with falling prices and short product cycles, a willingness to switch becomes a competitive advantage.
A step forward with sobering meaning
GPT-6 Sol and Luna are not proof that artificial intelligence is now error-free, autonomous, or universally applicable. The published success rates still reveal significant gaps. Even strong models regularly fail in demanding professional and technical tests. Therefore, these models do not replace the need for clear processes, good data, and responsible oversight.
Their economic significance lies elsewhere. OpenAI lowers the price of advanced AI so drastically that new applications become profitable and existing systems can be used more frequently. Sol brings powerful agentic and technical capabilities to a significantly lower price point. Luna makes large volumes of simple to medium tasks almost a mass-market infrastructure service.
The most important metric here is not the price per million tokens, nor is it a single benchmark value. What matters is the price of a reliably solved business problem. Sol and Luna appear to significantly improve this relationship, but the vendor data needs to be validated in real-world business environments. Those who simply buy cheaper tokens may end up spending more. In contrast, those who use models strategically, cache recurring context, and systematically measure quality can achieve a genuine leap in productivity.
Thus, Sol and Luna mark less the end of a technological race than the beginning of fiercer economic competition. Intelligence is becoming cheaper, but its successful use remains challenging. The advantage is shifting from companies with access to a single, robust model to those that can precisely orchestrate many models, understand their costs, and consistently manage their errors. This is precisely the true provocation of this publication: it is not artificial intelligence that is becoming scarce, but rather the ability to organize it in an economically viable way.
🎯🎯🎯 Data-driven B2B industry hub as a quasi-in-house solution
The quasi-in-house solution: How Xpert.Digital closes operational gaps in B2B marketing and sales – Smart Content-Driven Business - Image: Xpert.Digital
Xpert.Digital is a data-driven B2B industry hub led by Konrad Wolfenstein . The company acts as an external, quasi-in-house solution for industrial partners, closing operational gaps in marketing, content, and sales – without requiring additional resources on the client side.
More information here:
Your global marketing and business development partner
☑️ Our business language is English or German
☑️ NEW: Correspondence in your native language!
I and my team are happy to be available to you as your personal advisor.
You can contact me by filling out the contact form here wolfenstein@xpert.digital:or simply call me at +49 7348 4088 965. My email address is
I'm looking forward to our joint project.

