“Scientists have proved that AI will never become Einstein” is a stronger conclusion than the research behind it supports. It is easy to turn that into another premature business claim: neural networks can handle routine work but supposedly cannot even discuss strategy.

Both simplifications get in the way of useful AI. A model can suggest an unexpected idea, and a manager can be wrong about their own. The practical question is how to distinguish an interesting hypothesis from a persuasive story, test it and determine who makes the decision.

What LLMs Can’t Jump actually argues

Tom Zahavy, a Google DeepMind researcher, is listed in the official ICML 2026 program with Position: LLMs can't jump. It is an argued research position. Conference inclusion does not turn its conclusions into a proven theorem or the official position of an entire company. ICML 2026 · Position: LLMs can't jump · Tom Zahavy

The author distinguishes pattern finding, logical deduction and creating new explanatory premises. Using Einstein as a case study, he argues that current LLMs lack the mechanism for that last, abductive jump. Deriving a theory from supplied premises is discussed as a possibility, not a completed model experiment. Tom Zahavy · LLMs can’t jump · author-hosted position paper

“Not yet demonstrated” is weaker than “impossible.” Establishing impossibility for a class of systems requires precise definitions of that class, its permitted mechanisms and what counts as discovery. A historical example and conceptual argument can frame a valuable question without establishing the limits of every future AI system.

Deduction, induction and abduction in business

An LLM is a large language model: a system trained on data that generates a continuation based on context. A working product can add search, computation, memory, software tools and checks. Its usefulness cannot be assessed solely through a single chat response.

Three forms of reasoning help structure this discussion. Deduction yields a necessary conclusion from true premises through a valid inference. Induction generalizes observations. Abduction concerns explanatory hypotheses; the term is used both for generating them and for selecting the best explanation. Those meanings are not identical. Stanford Encyclopedia of Philosophy · Abduction

Three forms of reasoning: illustrative business examples
FormStructureExample
DeductionRule and specific case → consequenceDiscounts above 10% need approval. A 15% discount is requested, so approval is required.
InductionObservations → tentative patternIn the sample, repeat contact more often preceded a purchase. Test whether the pattern recurs.
AbductionUnexpected result → possible explanationCustomers postpone buying. A one-off payment may be inconvenient. The explanation needs testing.

The last row is not a proven diagnosis. Customers might be postponing purchases because of seasonality, a paused project or a changed budget. A hypothesis becomes valuable when it helps design a test that distinguishes these possibilities.

There is also a distinction between logic and model output. Valid deduction preserves the truth of premises. LLM-generated text that looks like a proof does not automatically have that guarantee. Recalculate numbers, open references and compare claims about the company with its actual data.

Einstein's jump: why change the starting question?

The free-fall thought experiment illustrates a change in how a phenomenon is described. In a small falling cabin, the observer and surrounding objects move together, producing local weightlessness. Gravity has not vanished everywhere: this concerns the chosen reference frame, with tidal effects limiting the approximation. The equivalence principle became one starting point for general relativity. Markus Pössel · Einstein Online · The equivalence principle

For our purposes, the useful technique is looking at a familiar situation from another position. “How can we make the salesperson call more often?” already assumes that call frequency limits performance. “Where does the customer lose interest?” allows more explanations and calls for different observations.

But reframing a business problem is not equivalent to inventing a new physical theory. Subscriptions instead of one-off sales or partners instead of direct selling are familiar options that AI can also suggest. They cannot demonstrate an exclusively human capacity for discovery.

Sometimes the difficult part is choosing the right problem. A useful formulation still requires its assumptions to be tested.

Can AI create something new? There are verifiable examples

“The model only repeats other people's work” is too strong. In AlphaEvolve, language models propose programs, automated evaluators test candidates and evolutionary search develops promising options. Google DeepMind reports new algorithmic results and improvements to computing infrastructure. This is a system for algorithm discovery and optimization with external evaluation of its outputs. Google DeepMind · AlphaEvolve: algorithm discovery and optimization

That does not necessarily mean inventing a fundamentally new language for physics. But dismissing the result as “combination” is equally unhelpful: a new working construction can have considerable scientific and practical value. Ask what was produced, how it was verified and which conditions people supplied.

For a business owner, the implication is practical. You do not need to settle the philosophical question of consciousness before using AI to search for solutions. You need to specify what would make a solution better than the previous one. If that criterion is wrong, a capable search process may efficiently improve the wrong outcome.

For example, a system might help increase the number of booked meetings. If the participants have neither need nor budget, calendars fill up without sales improving. The success criterion should account for downstream outcomes, contact quality and the team's effort.

Sales fell by 18%: what should the model do?

Imagine a company whose monthly revenue falls by 18%. This is an illustrative scenario, not an AI Office customer case. First, check the number itself: are the periods comparable, are returns treated consistently, has the export changed, or did a large one-off sale disappear?

Then break down the change: inquiry volume, qualified leads, conversion between stages, average order value and repeat purchases. Corporate AI with suitable integrations can prepare calculations, retrieve documents and gather observations. Access to a CRM does not mean access to every reason a customer declined.

Suppose “expensive” appears frequently in notes. Several explanations are possible: the price exceeds alternatives; the proposal fails to communicate value; customers need a different payment structure; or salespeople use “expensive” as a convenient label for every loss. The model can help assemble that list. Word frequency does not select the correct explanation.

The manager and team then decide which uncertainty to investigate first. A useful request is: “For each hypothesis, identify supporting evidence, contradictions, missing information and a way to test it.” “Reduce prices by 10%” without that chain turns a guess into an action too early.

An illustrative 18% revenue decline leads to three hypotheses: price, proposal value and payment format. Evidence testing precedes an accountable decision. aioffice.su.
Illustrative business scenario. AI and people can propose hypotheses together; an observation does not establish a cause, and a recommendation does not authorize action. Download image

Testing could combine a review of lost deals with a limited comparison of two proposal formats. Define the outcome in advance and account for group differences, seasonality and channel changes. A simple before-and-after sales comparison does not always establish that a new proposal caused the result.

Who chooses the objective when the model suggests a strategy?

The idea to sell a product as a service might come from a person or a model. Either way, questions follow: who finances equipment, who provides maintenance, how does cash flow change, and will customers accept recurring payments? Naming a business model does not answer them.

AI can prepare scenarios, check arithmetic and list weak points. Management needs to define constraints: acceptable investment, service quality, commitments to employees and risk tolerance. Different objectives produce different “best” decisions even with identical inputs.

Maximizing next month's revenue, retaining long-term customers and reducing employee overload are three different objectives. Sometimes they align; sometimes they require a trade-off. That trade-off should not silently become one convenient metric merely because it is easy to calculate.

There is no need to defend the human role with a claim of permanent intellectual uniqueness. A project needs clear authority: who sets the objective, who may change a budget, who accepts a proposal and who answers for the consequences. A more capable model makes that clarity more valuable.

AI Chief of Staff: preparing decisions without pretending to know everything

An AI Chief of Staff can be understood as an assistant for coordination and decision preparation. A useful version gathers context, highlights inconsistencies, tracks assignments and compares options. Permission to act is defined separately from the ability to write a recommendation.

A good AI management brief should contain more than a conclusion. It needs sources and data dates, calculation assumptions, alternative explanations, an expected effect and conditions that would trigger reconsideration. Missing information should remain visible.

Separate proposing, authorizing and executing. Routine actions can have predefined boundaries. Changes to prices, budgets, contractual terms or commitments to people need the company's chosen approval process. A persuasive model response does not expand its permissions.

World models: a possible next step, not a finished answer

Zahavy proposes linking reasoning to physically consistent world models that support action and observation. This is a proposed research direction, not a demonstrated solution to the limitation under discussion. Tom Zahavy · LLMs can’t jump · author-hosted position paper

A world model represents how an environment changes and responds to actions, allowing scenarios to be explored with feedback. A convincing simulation is not necessarily accurate, however. Google's Project Genie description explicitly warns that generated worlds may not follow real-world physics. Google · Project Genie: interactive worlds and limitations

Businesses encounter an analogous limit even in a spreadsheet. Sales after a price reduction can be calculated perfectly if increased demand is supplied as an assumption. But if that increase is invented, neat formulas do not make the forecast reliable. Scenario modeling and observing actual customers serve different purposes.

How to evaluate a management assistant in a pilot

Choose one recurring management question: project delays, inquiry quality or differences between planned and actual margins. Gather documents, identify the data owner and define what a useful outcome will look like before starting.

Test ordinary cases, exceptions and incomplete data. A good answer sometimes acknowledges insufficient information. If the assistant confidently explains every result even when a month is missing from the spreadsheet, that is a problem rather than evidence of deep analysis.

  • Check that calculations are reproducible and cited sources can be opened.
  • Look for a clear distinction between observations and assumptions.
  • Request alternative explanations and evidence that distinguishes them.
  • Measure preparation time through to an accepted decision, including review and corrections.
  • Ensure that recommendations do not lead to actions beyond authorized permissions.

What happens to the released time matters too. If a manager receives summaries faster but spends the rest of the time repairing incomplete data, the benefit is limited. If there is more time to speak with customers, test a different process and remove an actual constraint, automation provides practical value.

How this relates to AI Office

We propose building AI Office around a company's specific data and processes. A local AI station, knowledge search, document preparation and coordination workflows can form the basis of a management assistant. Features, integrations, answer quality and action permissions are defined and tested within the project.

Local deployment helps organize control over data. It does not make a model an infallible strategist. Decisions require current sources, transparent assumptions, verification and agreed authority regardless of where computation happens.

Our aim is to release management time for customers, employees and testing new directions. AI can participate in ideation, analysis and execution. Start with one process where its contribution can be evaluated on your company's real materials, then decide which responsibilities to delegate next.