A small computer sits on a desk. One employee asks it to shorten an email, another to find conditions in documents, and a third to explain internal software. The box looks the same, but very different models may be doing the work. Their selection helps determine whether this becomes a useful tool or an expensive experiment.

The new M6 and M5 Pro Mac mini went on sale on September 22, 2026. Apple announced them on August 25, with US starting prices of $899 and $1,699 respectively. Those are base configurations, not maximum-memory systems or Russian retail quotations. Apple: new Mac mini and Mac Studio available September 22 ↗ Apple: M6 and M5 Pro Mac mini announcement and US starting prices ↗

For AI Office, the opportunity is a compact local node for a defined company workflow. We examine five levels, from a small classifier to general-purpose Qwen3.8-27B and a large coding specialist. Information was checked on September 23, 2026. This is a shortlist for evaluation, not a report of our own testing on the new Mac mini.

Reviewers praise performance, but buying advice is mixed

“All reviewers call it the best buy” overstates the consensus. Macworld, for example, praises the M5 Pro model's size, quiet operation and performance but considers its higher price a substantial obstacle to recommending it. Macworld: Mac mini M5 Pro review ↗

A business owner faces a more specific question: what work will the machine do, and what will it cost to produce an accepted result? A good desktop Mac does not automatically make a good shared AI server. Video editing, short chats and an agent investigating a repository for half an hour impose different loads.

Start with tasks and memory. The newer chip number alone does not make M6 preferable to M5 Pro for every model.

M6 or M5 Pro: memory is the main dividing line

Apple lists 16, 24 and 32 GB unified-memory options for M6, and 24, 48 and 64 GB for M5 Pro. Memory bandwidth is 153–170 GB/s for M6 depending on configuration, and 307 GB/s for M5 Pro. Apple Mac mini technical specifications ↗

Unified memory is shared by the system and compute components. The model does not receive a separate unlimited pool of video memory: macOS, the browser, knowledge retrieval and other applications need room too. Select memory at ordering time around the whole workload.

Bandwidth describes data transfer capacity, not a ready-made words-per-second figure. Actual performance depends on model architecture, weight format, runtime, context and concurrent requests. Twice the bandwidth does not guarantee twice the workflow speed.

How to read the recommendations

The table uses specific four-bit MLX Community conversions. These are community-converted weights, not separate Apple models. Weight sizes are sums of published safetensors files, rounded in decimal gigabytes. Tokenisers and other supporting files are excluded.

Runtime memory adds context state, intermediate computation, visual inputs and overhead. Some system memory may also be unavailable to the selected graphics backend. Whether weights fit and whether the complete system operates reliably are separate tests.

The table is our preliminary engineering assessment for one loaded model and one active request. An initial trial can use roughly 4,000–8,000 context tokens, increasing the limit while measuring memory. These are proposed pilot conditions, not measured minimum requirements or a guarantee of operation.

Four-bit MLX pilot candidates. Weight size is not runtime memory; we have not tested these configurations
ModelWeights onlyProposed taskConfiguration to evaluate
Qwen3.5-2B1.72 GBCategories, fields, brief summariesM6 · 16 GB
Qwen3.5-4B3.03 GBEmails and short draftsM6 · 16–24 GB
Qwen3.5-9B5.95 GBAnswers grounded in retrieved documentsM6 · 24–32 GB
Qwen3.8-27B16.05 GBDocument comparison and complex tasks32 GB for a trial; M5 Pro · 48–64 GB for node evaluation
Qwen3-Coder-Next44.84 GBSpecialised coding agentM5 Pro · 64 GB, separate experiment

B denotes billions of parameters. More parameters do not guarantee universal superiority: specialisation, training and the particular conversion matter. Original-model benchmark results should not automatically be assigned to a quantised version.

Level 1. Qwen3.5-2B for narrow, well-defined operations

The official Qwen3.5-2B card primarily positions it for prototyping and task-specific development. A four-bit MLX conversion has approximately 1.72 GB of weights. Qwen3.5-2B: official model card ↗ mlx-community/Qwen3.5-2B-4bit weight files ↗

Candidate tasks include assigning short requests to a few categories, extracting fields from clean text, normalising names and producing brief summaries. Clear rules and easily verified output make a small model more useful.

For example, a request says “we need another delivery of filters”. The model suggests the procurement category and extracts the product name. The application validates the category while leaving a missing product identifier empty. Inventing an identifier merely to complete the form would fail acceptance.

An M6 with 16 GB is a starting candidate for this pilot. Buying M5 Pro solely for such a classifier would need justification from another workload or demonstrated bottleneck. A complex contract full of exceptions and cross-references is a poor starting task for the smallest model.

Level 2. Qwen3.5-4B for emails and short drafts

Qwen3.5-4B is the next candidate for everyday writing; the selected four-bit MLX conversion contains approximately 3.03 GB of weights. The original model card and converted files are available separately. Qwen3.5-4B: official model card ↗ mlx-community/Qwen3.5-4B-4bit weight files ↗

Proposed tasks include shortening an email without losing commitments, turning notes into prose, drafting alternative replies and summarising a short document. Specify the tone, audience and facts that must remain unchanged.

Review is concrete: are dates, amounts, names and obligations preserved? A smoother email that changes “we plan to ship” into “we guarantee shipment” fails acceptance. A person approves the text before sending it.

An M6 with 16 GB is a candidate for a bounded trial; 24 GB provides more headroom for ordinary applications. This level suits initial exploration before a company builds complex tool workflows.

Level 3. Qwen3.5-9B for internal knowledge assistance

Include Qwen3.5-9B when evaluating more substantial answers. Its selected four-bit MLX weights total approximately 5.95 GB. Qwen3.5-9B: official model card ↗ mlx-community/Qwen3.5-9B-4bit weight files ↗

A useful pilot involves questions about instructions, service catalogues and internal procedures. Retrieval first selects relevant passages; the model then answers from them. This is commonly called RAG: generation is grounded in retrieved material rather than training knowledge alone.

If an employee asks which documents are needed for a building handover, the answer should provide a current list and supporting references. Conflicting procedures should be exposed rather than silently resolved in favour of a convenient version.

For procurement around this workflow, we would first evaluate M6 with 24–32 GB. That provides room for documents, retrieval and the application; it does not mean a 9B model categorically cannot run on 16 GB. Long conversations and concurrent employees need a separate load test.

Level 4. Qwen3.8-27B as the main general-purpose candidate

Qwen3.8-27B combines a 27-billion-parameter language component with a vision encoder. Its developer positions it for coding, professional work and multistep tasks. The selected four-bit MLX conversion contains approximately 16.05 GB of weights. Qwen3.8-27B: official model card ↗ mlx-community/Qwen3.8-27B-4bit weight files ↗

Within this shortlist, it is our main general-purpose candidate for a more demanding corporate pilot: compare supplier proposals, identify differences in conditions, draft an internal memo or interpret an ambiguous request before preparing a quotation.

The useful sequence is to find evidence, identify missing information, develop a conclusion and propose a next step. Verifiable application code should calculate amounts, and actions should remain within explicit permissions. A larger model should not become the only reviewer of its own mistakes.

M6 with 32 GB is a candidate for short-context, sequential testing. For a dedicated node with additional services, we would start evaluation at M5 Pro with 48 GB; 64 GB adds headroom. Even 64 GB does not promise the full context advertised in the model card with arbitrary concurrency.

This is the upper general-purpose tier of our selection, not a claim to the world's best model. Russian-language quality, reference accuracy and tool use need comparison on your documents. We have not measured them on the new Mac mini.

Level 5. Qwen3-Coder-Next for demanding development work

A company with its own software has a separate candidate: Qwen3-Coder-Next. Its official card describes a coding-agent model with 80 billion total parameters and three billion active per step. The selected four-bit MLX conversion contains approximately 44.84 GB of weights alone. Qwen3-Coder-Next: official architecture and intended use ↗ mlx-community/Qwen3-Coder-Next-4bit weight files ↗

“Only 3B active” does not mean only three billion parameters must be stored. The architecture selects part of the computation while other weights must remain available. A large mixture-of-experts model can therefore need substantial memory despite relatively few active parameters.

Of the new Mac mini options, M5 Pro with 64 GB is the relevant candidate for this experiment. It is resource-constrained, not a guaranteed comfortable configuration. Check backend memory availability, short context and a single request without other large models loaded. Persistent swapping or allocation failures mean the configuration fails acceptance.

Proposed work includes explaining a repository, preparing changes, tests and migrations, and investigating errors. The agent needs an isolated working copy and limited permissions. Changes pass tests and review; loading a large model does not justify letting it release code without control.

This is the largest model listed and a development specialist. It is not automatically better than Qwen3.8-27B for emails, procurement or management summaries. If dependable general assistance is the priority, leaving memory free may be more useful than filling it with the largest possible weights.

Runtimes: downloading weights is only one step

MLX LM runs language models on Apple silicon; MLX-VLM serves models that combine image and text processing. The right path depends on the architecture and conversion. MLX LM: language models on Apple silicon ↗ MLX-VLM: vision-language inference on Mac ↗

Distinguish original weights, MLX conversions and other formats such as GGUF. A file for one runtime should not be assumed compatible with another. Record the exact model identifier, file revision, runtime version and context settings so a pilot result can be reproduced.

Vision support also does not create a complete service for arbitrary scanned documents. PDF intake, text extraction, page processing and field verification need their own workflow. Installing a vision model does not automatically add OCR to AI Office.

Local inference can allow document processing without sending requests to an external provider. Verify the entire chain, including the interface, retrieval, logs and tools. Model downloads and updates are usually separate operations; optional cloud features do not become local because a model is installed.

Choosing a configuration for the work

For emails, short summaries and narrow operations, begin comparisons with M6 at 16–24 GB and 2B–4B models. More expensive hardware needs justification from a confirmed constraint or other desktop work.

For knowledge assistance, evaluate M6 at 24–32 GB with a 9B model. If documents are complex, add Qwen3.8-27B on 32 GB to the trial. Compare correctness, waiting time, memory and human rework together.

For a dedicated corporate node using Qwen3.8-27B, consider M5 Pro at 48–64 GB. Purchase should follow validation of the selected workflow, particularly for multiple employees. Memory headroom helps but does not turn one node into an unlimited server.

Qwen3-Coder-Next needs a separate experiment at 64 GB. If the requirement combines a heavy model, large context and several active agents, consider hardware with more memory headroom. A compact computer has a practical ceiling.

Size storage for several model versions, documents and backups. An external SSD can expand storage but does not replace working memory for fast inference. Disk capacity and runtime memory solve different problems.

What to measure before buying several machines

Prepare 20–30 representative tasks, including short, long and ambiguous cases, plus questions where the correct response is that information is missing. This is a suggested pilot set, not an official standard. Define acceptance criteria in advance.

Measure time to the first response and full completion, memory, swapping, corrections and employee review cost. Repeat under simultaneous requests and background document indexing. That workload will show whether the machine suits a shared service.

Do not equate installed model count with concurrent assistants. Many models can sit on disk, but simultaneous loading consumes additional memory. A queue and sequential execution may be more economical than trying to keep everything active.

How Mac mini could fit into AI Office

We view it as a possible hardware basis for a local station that needs scenario-specific validation. The current AI Office prototype includes document retrieval with permissions and sources, Excel/CSV catalogues, verifiable quotation calculations and creation of a local task after human approval.

Ollama and a compatible HTTP provider are supported. That is an integration starting point, not confirmation that every MLX conversion listed is compatible. We have not tested the M6 or M5 Pro Mac mini, model speed or shared AI Office workloads on them. This article does not announce a finished OCR service or coding agent either.

The value lies in combining hardware, a suitable model and a working process. First, the computer helps complete one task with a verifiable result. The company can then decide which capabilities are worth expanding.

A conversation with AI Office can start with one question: what should your local AI do every day? Choose the model and memory around that answer, so the small box becomes part of the work rather than a collection of downloaded models.