Which graphics card should I buy for local AI?

Choose what you want to do, compare a few cards, then check the exact listing. For desktop upgrades in the US. Prices cover the graphics card only.

1. What would you like to do?
2. Narrow the recorded offers (optional)

Amazon US · USD · Recorded offers, not live checkout prices. Each quote has its own date. Fresh quotes are preferred; older quotes are clearly identified. New and Used prices can come from different dates. As an Amazon Associate I earn from qualifying purchases.

Try it first

Use your existing computer

You do not need to buy a graphics card just to install a local AI app. Try a small model first; your system still needs enough memory, and CPU generation can be slower.

Start with Ollama and a small Qwen3 4B model. On a compatible computer with enough system memory, CPU or supported integrated/unified-memory hardware may be enough to explore. You can decide whether to upgrade after trying it.

Everyday chat

Start by comparing 12–16GB cards

A starting point for small, compressed chat models, such as an 8B model at 4-bit precision. More memory leaves room for the conversation and the app.

Qwen3 8B Q4_K_M is a 5.2GB download. That is the model file, not its total working memory. Model package sizes.

These are planning ranges, not guaranteed fit or speed. Adjust the 8B memory example for your model and conversation length.

RTX 3060 · 12GB

A 12GB starting option. Confirm the 12GB version, rather than the 8GB version.

$399.99 · Used Last recorded quote

Quote dated 2026-10-05 08:57:27 UTC. Older quote; confirm today’s price and stock before deciding.

Check used price on Amazon for RTX 3060 12GB
Exact listing

MSI Gaming GeForce RTX 3060 12GB 15 Gbps GDRR6 192-Bit HDMI/DP PCIe 4 Torx Triple Fan Ampere OC Graphics Card (RTX 3060 Ventus 3X 12G OC) (Renewed)

$479.99 · New Last recorded quote

Quote dated 2026-10-02 09:42:54 UTC. Older quote; confirm today’s price and stock before deciding.

Check new price on Amazon for RTX 3060 12GB
Exact listing

ZOTAC Gaming GeForce RTX 3060 Twin Edge 12GB GDDR6 Gaming Graphics Card

Compare all RTX 3060 listings

RTX 4060 Ti · 16GB

A 16GB option. This shortlist uses the 16GB version, not the 8GB version.

$699.97 · Used Last recorded quote

Quote dated 2026-09-29 16:18:54 UTC. Older quote; confirm today’s price and stock before deciding.

Check used price on Amazon for RTX 4060 Ti 16GB
Exact listing

ZOTAC Gaming GeForce RTX 4060 Ti 16GB AMP DLSS 3 16GB GDDR6 128-bit 18 Gbps PCIE 4.0 Compact Gaming Graphics Card, IceStorm 2.0 Advanced Cooling, Spectra RGB Lighting, ZT-D40620F-10M (Renewed)

$949.99 · New Last recorded quote

Quote dated 2026-09-29 16:17:45 UTC. Older quote; confirm today’s price and stock before deciding.

Check new price on Amazon for RTX 4060 Ti 16GB
Exact listing

MSI GeForce RTX 4060 Ti Ventus 3X 16G OC Graphics Card -NVIDIA RTX 4060 Ti, 16GB GDDR6 Memory, 18Gbps, PCIe 4.0, DLSS3

Compare all RTX 4060 Ti listings

RTX 5060 Ti · 16GB

Another 16GB option. Compare today’s price with the 4060 Ti; this guide does not rank their speed.

$788.50 · New Last recorded quote

Quote dated 2026-09-29 14:05:39 UTC. Older quote; confirm today’s price and stock before deciding.

Check new price on Amazon for RTX 5060 Ti 16GB
Exact listing

ASUS Prime GeForce RTX 5060 Ti 16GB GDDR7 Gaming Graphics Card

Compare all RTX 5060 Ti listings

Coding & more room

Compare 16GB and 24GB cards

More room to experiment with a 14B compressed model or keep more conversation in memory. Coding is also possible with smaller models; long documents and coding agents can need much more memory.

Qwen3 14B Q4_K_M is a 9.3GB download, before conversation memory and runtime overhead. Model package sizes.

These are planning ranges, not guaranteed fit or speed. Adjust the 14B memory example for your model and conversation length.

RTX 5060 Ti · 16GB

Another 16GB option. Compare today’s price with the 4060 Ti; this guide does not rank their speed.

$788.50 · New Last recorded quote

Quote dated 2026-09-29 14:05:39 UTC. Older quote; confirm today’s price and stock before deciding.

Check new price on Amazon for RTX 5060 Ti 16GB
Exact listing

ASUS Prime GeForce RTX 5060 Ti 16GB GDDR7 Gaming Graphics Card

Compare all RTX 5060 Ti listings

RTX 5070 Ti · 16GB

16GB of model-memory room, just like the other 16GB options. Paying more does not add VRAM.

$1,185.99 · New Last recorded quote

Quote dated 2026-10-02 09:42:54 UTC. Older quote; confirm today’s price and stock before deciding.

Check new price on Amazon for RTX 5070 Ti 16GB
Exact listing

PNY NVIDIA GeForce RTX™ 5070 Ti OC Triple-Fan Graphics Card

Compare all RTX 5070 Ti listings

RTX 3090 · 24GB

A 24GB option, including recorded Used offers. Check power, case space and return terms.

$2,195.00 · New Last recorded quote

Quote dated 2026-10-07 06:17:19 UTC. Older quote; confirm today’s price and stock before deciding.

Check new price on Amazon for RTX 3090 24GB
Exact listing

nVidia GeForce RTX 3090 Founders Edition Graphics Card

$1,899.99 · Used Last recorded quote

Quote dated 2026-10-05 08:51:20 UTC. Older quote; confirm today’s price and stock before deciding.

Check used price on Amazon for RTX 3090 24GB
Exact listing

NVIDIA GeForce RTX 3090 Founders Edition Graphics Card (Renewed)

Compare all RTX 3090 listings

Larger models

Compare 24GB and 32GB cards

A starting point for experimenting with compressed 30–32B models. A 24GB card has less room for conversation memory than a 32GB card; check the exact model and context before buying.

Qwen3 32B Q4_K_M is a 20GB download. Its 8-bit version is 35GB before overhead, so the two versions need different hardware. Model package sizes.

These are planning ranges, not guaranteed fit or speed. Adjust the 32B memory example for your model and conversation length.

RTX 3090 · 24GB

A 24GB option, including recorded Used offers. Check power, case space and return terms.

$2,195.00 · New Last recorded quote

Quote dated 2026-10-07 06:17:19 UTC. Older quote; confirm today’s price and stock before deciding.

Check new price on Amazon for RTX 3090 24GB
Exact listing

nVidia GeForce RTX 3090 Founders Edition Graphics Card

$1,899.99 · Used Last recorded quote

Quote dated 2026-10-05 08:51:20 UTC. Older quote; confirm today’s price and stock before deciding.

Check used price on Amazon for RTX 3090 24GB
Exact listing

NVIDIA GeForce RTX 3090 Founders Edition Graphics Card (Renewed)

Compare all RTX 3090 listings

RTX 4090 · 24GB

24GB, the same capacity as the 3090. A higher price does not add model-memory room; speed is a separate question.

$4,450.00 · New Last recorded quote

Quote dated 2026-10-02 09:42:54 UTC. Older quote; confirm today’s price and stock before deciding.

Check new price on Amazon for RTX 4090 24GB
Exact listing

MSI RTX 4090 Gaming Trio 24G Carte graphique NVIDIA GeForce RTX 4090 24 Go GDDR6X

$3,199.99 · Used Last recorded quote

Quote dated 2026-09-29 16:21:21 UTC. Older quote; confirm today’s price and stock before deciding.

Check used price on Amazon for RTX 4090 24GB
Exact listing

PNY GeForce RTX™ 4090 24GB Verto Triple Fan Graphics Card DLSS 3 (Renewed)

Compare all RTX 4090 listings

RTX 5090 · 32GB

32GB gives more memory room than 24GB. It is a costly option, and does not guarantee that every large model fits.

$8,099.00 · New Retrieved within 24 hours

Quote dated 2026-10-08 21:25:58 UTC. Amazon’s page date is unknown; retrieval does not prove when Amazon set this price.

Check new price on Amazon for RTX 5090 32GB
Exact listing

ZOTAC Gaming Geforce RTX 5090 Amp Extreme Infinity Nvidia 32 Gb, W129163525 (Extreme Infinity Nvidia 32 Gb Gddr7)

Compare all RTX 5090 listings

The terms, in plain English

VRAM
The card’s memory. More room helps hold the model and conversation; it does not guarantee a faster or smarter model.
8B / 14B / 32B
The approximate model size in billions of parameters. Larger models usually need more memory, but size alone does not measure quality.
4-bit / Q4
A compressed version of model weights that uses less memory, with possible quality tradeoffs. Compare the exact model version.
Context
How much text the model can keep track of at once. Longer chats, documents and multiple users can need more memory.

Before you buy

  1. Confirm the exact memory version. RTX 4060 Ti and RTX 5060 Ti come in both 8GB and 16GB versions.
  2. Check your case space, power supply and power connectors against the exact card manufacturer’s specifications. This is a card upgrade, not a complete computer.
  3. Check your chosen app’s GPU, operating-system and driver support. This shortlist uses NVIDIA’s documented CUDA path; AMD, Intel and Apple can also work with supported backends.
  4. For Used cards, inspect the listing condition and return terms. We have not tested these individual cards.
  5. Check Amazon’s current price and availability. A recorded quote is not a promise of today’s price or a claim of the best deal.
What about 70B models or two GPUs?

A dense 70B model at ideal 4-bit precision needs about 32.6GiB for weights alone, before compression metadata, conversation memory and app overhead. A 32GB card is not a safe all-on-GPU recommendation for that case. Multi-GPU setups and CPU offloading depend on the runtime and are beyond this beginner single-card shortlist. Use the VRAM calculator to plan allowances.

Why these cards and categories?

These are editorial starting choices, selected for documented NVIDIA runtime support and exact memory variants. They are not benchmark rankings, Amazon best-seller recommendations or claims to outperform every alternative. We have not measured tokens per second. Cards requiring external enclosures, special BTF motherboards or liquid-cooling setups are excluded from the shortlist.

For each exact card and condition, we show the lowest qualifying fresh quote when available. Otherwise, we show the lowest recorded quote from the most recent quote date in our 30-day sample. These are selected catalog offers, not market-wide lowest prices.

The memory bands are our planning interpretation of published model package sizes plus runtime/context overhead. The calculator’s 2GiB conversation-cache, 1GiB runtime and 20% headroom examples are adjustable illustrations, not measured requirements.

Sources and next steps

Ollama GPU support · Ollama loading and CPU/GPU FAQ · Context and memory · Qwen3 model formats · llama.cpp supported backends · NVIDIA memory variants

Want to compare more? Full GPU price table · Fresh price-by-memory sample · Detailed VRAM calculator.

Maintained by Triftan. Guide reviewed 2026-10-10; quote dates are separate. Affiliate disclosure.