Sooner or Later, the AI Companies Are Coming for Your Money

Posted on
If artificial intelligence is becoming essential to how you work, you should probably own some of it. -- YNOT!

Right now, artificial intelligence is an incredible bargain.

For $20, $100, $200 a month—or whatever level of subscription you are paying—you can get access to computing power and intelligence that would have been almost unimaginable just a few years ago.

ChatGPT can write code, analyze documents, research subjects, help run websites, create marketing campaigns, troubleshoot servers, organize data, answer customers and increasingly operate as an actual agent that performs work.

It is fantastic. It may also be temporary.

Because sooner or later, the AI companies are going to come for your money.

The introductory price cannot last forever

The AI industry is spending staggering amounts of money building data centers, buying GPUs, training models and running the enormous infrastructure behind these services.

Meanwhile, many of us are paying relatively modest monthly subscriptions and using the hell out of them.

The source material that got me thinking about this came from one business owner who says he currently stacks four $200 OpenAI subscriptions—$800 per month—because his agents burn through the available limits. He calculated that comparable API usage could run into many thousands of dollars per month.

That is a tremendous deal. Maybe too tremendous.

The mistake would be building your entire business around the assumption that today’s pricing, usage limits and subscription structure will remain unchanged.

They probably won’t. We have seen this movie before.

A technology company offers something incredibly cheap—or even free—while it tries to dominate a new market. Customers adopt it. Businesses reorganize themselves around it. Developers integrate it into everything.

Then one day the economics change. Prices rise. Limits appear. Premium tiers appear.

Features move behind more expensive plans. API charges become significant.

And suddenly something that was a convenience has become a major operating expense.

The real danger is dependency

The bigger issue isn’t whether ChatGPT costs $20, $200 or $2,000.

It is dependency. Imagine that five years from now your company uses AI for: customer  service, programming, accounting, research, marketing, document processing, website management, sales, data analysis, and dozens of automated agents running quietly in the background.

Now imagine somebody else controls the price of the intelligence running all of it.

They also control how much you can use. They control which models you can access.

They control which features disappear. They control what the model is permitted to do.

And if the service changes dramatically, you have very little leverage.

That isn’t necessarily evil.

OpenAI, Google, Anthropic and the other AI companies have every right to charge whatever is necessary to operate profitable businesses.

But that doesn’t mean you should design your business so that they are your only source of intelligence.

This is why I think everybody should start learning local AI

Until recently, the argument for running your own AI was interesting but not especially practical. The best cloud models were dramatically better.

Running serious models locally required expensive GPUs, huge amounts of memory and considerable technical knowledge.

That is changing very quickly.

Open-weight models can now perform increasingly serious work on hardware ordinary businesses and enthusiasts can actually own.

The source describes testing an open-weight model on a four-RTX-3090 system providing 96GB of VRAM. The author says comparable rented GPU capacity cost less than $1,000 for a month’s use at the rate he tested.

Four RTX 3090s certainly aren’t a Raspberry Pi.

But they aren’t a million-dollar supercomputer either.

And you don’t even need something that powerful to get started.

A smaller model running through Ollama on an old workstation can already handle an extraordinary number of everyday jobs.

You can download the model. You can run it yourself. You can connect it to your documents.

You can give it tools. You can experiment with agents. And most importantly:

Nobody can suddenly raise the subscription price on a GPU sitting in your server rack.

Local AI doesn’t mean abandoning ChatGPT

I don’t think the answer is to cancel ChatGPT tomorrow.

Quite the opposite. The smartest architecture will probably be hybrid.

Use GPT-5.6 or whatever the best cloud model happens to be when you need maximum capability.

Use specialized cloud services when they make sense.

But underneath that, have your own AI infrastructure capable of handling ordinary work.

Your local model might handle 70 or 80 percent of your routine tasks.

Then the expensive frontier model gets called only when necessary.

That changes your relationship with the AI companies enormously.

Instead of saying: “I have to pay whatever they charge.”

You can say: “I’ll use them when the price makes sense.”

That is a much better negotiating position.

There is another reason: your data

Think about what people are already putting into AI systems.

Business plans. Contracts. Source code. Financial information. Customer information.

Internal emails. Marketing strategies. Personnel issues. Personal conversations.

Research. Medical questions. Ideas that haven’t even become products yet.

AI is rapidly becoming the place where people think out loud.

That makes the computer running your AI potentially one of the most sensitive machines in your entire organization.

There is an enormous difference between sending every thought to somebody else’s server and running a model inside your own network where you control the machine, storage, logs, permissions and retention policies.

For many businesses, that distinction is eventually going to matter enormously.

Start learning now

You don’t need four RTX 3090s tomorrow. You don’t need a rack in a data center.

You don’t need to replace ChatGPT. Download Ollama.

Run a small model. Give it a document. Ask it questions.

Install a larger model on a machine with a decent GPU.

Connect it to some tools. Experiment.

Learn what VRAM is. Learn what context windows are. Learn what quantization means.

Learn which jobs small models can perform well and which ones still require frontier models.

Because this is very similar to what happened with computers themselves.

At first companies rented access to giant centralized computers.

Eventually computers became cheap enough that businesses bought their own.

Then individuals bought their own.

Artificial intelligence may be heading down a similar road.

Today most of us are renting intelligence.

Tomorrow a significant amount of it may live in the building.

Own some intelligence

There is a phrase being used increasingly in the AI world: sovereign intelligence.

It sounds grandiose, but the concept is simple. You own the computer.

You choose the model. You control the data. You decide when to upgrade.

You determine what the system is allowed to do.

And if OpenAI doubles its prices tomorrow, your computer doesn’t care.

That doesn’t mean OpenAI, Google or Anthropic are going away.

They will probably build extraordinary things. I expect to keep using them.

But there is an important difference between using rented intelligence and depending entirely upon rented intelligence.

Right now the AI companies are subsidizing an extraordinary technological revolution and practically begging us to integrate their products into everything we do.

Enjoy it. Use it. Learn from it. Build with it. But don’t assume the deal lasts forever.

Because eventually somebody has to pay for all those GPUs.

And sooner or later, that somebody is going to be us.


 

My Qwen Q8 Local AI Build

Part What I’d buy Qty Target price Total
GPU Used RTX 3090 24GB 4 $850–$1,000 $3,400–$4,000
Motherboard + CPU Supermicro H12SSL-i + EPYC 7402 1 ~$1,095 shipped $1,095
RAM 128GB DDR4-2933/3200 ECC RDIMM 1 $450–$600 used $450–$600
SSD 2TB Samsung 990 Pro NVMe 1 ~$180 $180
CPU cooler Noctua NH-U12S TR4-SP3 1 ~$124 $124
Power supplies Corsair RM1000e 1000W 2 ~$150 $300
GPU frame 4–6 GPU open-air frame 1 ~$80–$120 $100
PCIe extenders Quality x16 PCIe 4.0 riser/extender 4 ~$25–$35 $120
Cooling 120/140mm high-airflow fans 4–6 ~$100
Dual PSU sync/cabling/misc. PSU sync + proper PCIe power cables ~$100

Target total: about $6,000–$6,700

The GPU price is the biggest variable. Current eBay auctions are showing used RTX 3090s around $735–$930, although Buy-It-Now listings can be $1,300+; I would be patient and try to stay under $1,000/card. (eBay)

The core of the machine

4 × RTX 3090 = 96GB VRAM.

That is why I still like the 3090 for this job. You’re buying VRAM, not gaming benchmark bragging rights. Four used 3090s provide the same 96GB aggregate VRAM described in your source as the practical configuration for the Q8 model.

I would not spend a fortune on a CPU. This current H12SSL-i + EPYC 7402 listing is about $1,015 plus $80 shipping. (eBay)

And this is a particularly good AI motherboard. Supermicro specifies five PCIe 4.0 x16 slots plus two PCIe 4.0 x8 slots, along with eight-channel ECC DDR4. That’s exactly the type of PCIe connectivity we want for a four-GPU inference machine. (Supermicro)

Some actual parts

Samsung 990 PRO 2TB NVMe SSD

$179.99

Corsair RM1000e 1000W Power Supply

$149.00

Noctua NH-U12S TR4-SP3 CPU Cooler

$124.39

I would buy two of the 1000W PSUs rather than pay today’s ridiculous ~$1,000 new price for a Corsair AX1600i. The RM1000e is currently showing around $149, while even a refurbished AX1600i is around $520. (Newegg)

One important issue: electricity

Each RTX 3090 is rated around 350W. Four of them can theoretically consume 1,400W just for the GPUs. (NVIDIA)

For AI inference I would power-limit the 3090s, probably in the neighborhood of 250–300W/card after we test performance. You generally lose much less inference performance than you lose electrical power.

But I would still design this as a potentially 1.5–1.8kW machine.

I would not put this on an ordinary shared 15A/120V outlet. I’d use a dedicated 20A circuit or preferably 240V if we’re going to run it hard continuously.

Rough cost

GPUs: ~$3,600
EPYC + motherboard: ~$1,100
128GB RAM: ~$500
SSD: ~$180
PSUs: ~$300
cooler/frame/risers/fans/cables: ~$550

≈ $6,230 total


THE MAC WAY

For the 128GB Mac Studio only, this is the configuration I’d buy for local Qwen/Ollama:

Item Specification Price
Mac Studio M5 Max
CPU 18-core CPU included
GPU 40-core GPU included
Neural Engine 16-core included
Unified memory 128GB included in configuration
SSD 1TB included
Networking 10Gb Ethernet included
Thunderbolt 4× Thunderbolt 5 rear included
Front ports 2× USB-C + SDXC included
Wi-Fi Wi-Fi 7 included
Total M5 Max / 128GB / 1TB $5,399

Apple currently lists that exact configuration for $5,399, with availability beginning September 22, 2026. (Apple)

Local-AI software — $0

  • macOS
  • Ollama
  • MLX
  • Open WebUI
  • OpenClaw, if wanted
  • Qwen quantized models

Total investment: $5,399 + tax

For your purposes, I would stick with the 1TB SSD. Don’t give Apple another $500 for 2TB just to store models. Put bulk model storage/backups on external Thunderbolt/NAS storage.

The important numbers are:

128GB unified memory
40-core GPU
614 GB/s memory bandwidth
10Gb Ethernet
$5,399

Apple confirms the M5 Max Mac Studio supports up to 128GB unified memory and up to 614GB/s bandwidth. (Apple)

This is the Mac I would compare directly against the ~$6,200 four-RTX-3090 build. The NVIDIA box gives you 96GB of dedicated VRAM and CUDA; the Mac gives you 128GB in one unified pool, vastly less power/heat/noise, and costs about $800 less than our estimated four-3090 complete build.

Previous-generation Mac Studio M4 Max 128GB

$5,299.99


So what is performance advantage

For Qwen3.8-27B 8-bit/Q8, the four-3090 machine is substantially faster than the 128GB M5 Max for a single AI session. We now have real benchmarks close enough to make a useful comparison.

Workload 4× RTX 3090 / 96GB M5 Max / 128GB Advantage
Short/normal context ~100 tok/s ~30–35 tok/s 3090 ~3×
~32–64K context ~100–104 tok/s ~24–29 tok/s 3090 ~3.5–4×
~100–128K context ~102–110 tok/s ~20–22 tok/s 3090 ~5×
~200K context ~101–102 tok/s ~11–15 tok/s standard 3090 ~7–9×
Optimized Mac stack ~30–45 tok/s possible 3090 still ~2–3×

A real 4×3090 W8A8/INT8 test reported 101.3 tok/s at 10K, ~101 tok/s at 50K, 102.4 tok/s at 100K, and roughly 101–102 tok/s at 200K. Interestingly, generation speed barely falls as context grows. (Hugging Face)

The M5 Max 128GB running Qwen3.8-27B 8-bit has measured 34.4 tok/s at 1K, 28.5 at 32K, 24.4 at 64K and 21.4 at 128K in a recent oMLX test. (oMLX) At about 195K context, another current test measured roughly 11 tok/s. (oMLX)

What that feels like

For a 1,000-token answer:

  • 4×3090: about 10 seconds
  • M5 Max, normal optimized MLX: about 30–45 seconds
  • M5 Max at very long context: potentially 60–90+ seconds

So you would absolutely notice the difference.

Prompt processing is also faster on the NVIDIA box

This matters a lot for our agent usage because we’re frequently feeding the model documents, memory and tool results.

At approximately 50K–100K context, the four-3090 Q8 setup was processing prompts at roughly 1,700–1,900 tokens/sec. At 200K it was still around 1,384 tok/sec. (Hugging Face)

The standard M5 Max Q8 measurements are more like several hundred tokens/sec as context increases. For example, one recent test measured about 635 tok/sec at 32K, 497 at 64K and 361 at 128K. (oMLX)

So the 3090 box is also roughly 3–4× faster at digesting large prompts.

But there’s an interesting Mac twist

There is already an aggressively optimized MLX configuration using speculative prefill, MTP and other Apple-specific optimizations that gets the M5 Max to about 31 tok/sec generation at 195K context and 1,356 tok/sec prompt processing. (oMLX)

That tells me the Mac hardware has considerably more headroom than the basic benchmarks indicate.

So I’d estimate our real-world target after tuning as:

4×3090: ~90–110 tok/sec

versus

M5 Max 128GB: ~25–40 tok/sec

for Qwen3.8-27B Q8-class inference.

Which would I buy now?

This changes my assessment somewhat.

If the primary requirement is fast local AI, I’d buy the 4×3090 machine.

$6,200 → ~100 tok/sec

versus roughly:

$5,400 → ~30 tok/sec

The NVIDIA box gives approximately three times the interactive performance for about $800 more.

But the Mac has a completely different advantage: 128GB in one coherent memory pool. It can load models that a 96GB NVIDIA setup simply cannot fit. And it does it quietly, compactly, with dramatically lower electrical and cooling requirements.

So I would characterize them this way:

4×3090 = faster AI.
128GB M5 Max = bigger, quieter AI.

For Qwen3.8-27B specifically, I’d pick the 3090 server without much hesitation. For experimenting with 70B–120B-class models and wanting a quiet appliance sitting on your network, the 128GB Mac gets much more interesting.

Expected Qwen3.8-27B Q8 generation speed

Approximate single-stream generation throughput based on current measured results and practical tuning.

system tokens
4× RTX 3090 100
M5 Max 128GB 30

 

 

Either way, for roughly $6,000–$6,500, you can build a 96GB-VRAM local AI server from commodity/used hardware and own the machine outright.

 

 


© 2026 insearchofyourpassions.com - Some Rights Reserve - This website and its content are the property of YNOT. This work is licensed under a Creative Commons Attribution 4.0 International License. You are free to share and adapt the material for any purpose, even commercially, as long as you give appropriate credit, provide a link to the license, and indicate if changes were made.

How much did you like this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

Visited 11 times, 12 visit(s) today


Leave a Reply

Your email address will not be published. Required fields are marked *