You want to run Ollama but you have an Intel Arc graphics card? I have some great news for you. Introducing IPEX-LLM for all your inference needs. You will be chatting with silicon in no time. This video will walk you how to use the Ollama Portable Zip on Intel GPU with IPEX-LLM. Enjoy!
0:00 Intro (Intel Arc Fans Rejoice)
0:55 Getting Started with IPEX-LLM
3:07 DEMO – Run Ollama on Intel Arc
5:43 DEMO – Initial Tests
11:03 DEMO – Testing Immediate Command Lists
12:08 DEMO – Compare GPU and CPU Performance
Run Ollama Portable Zip on Intel GPU with IPEX-LLM
https://github.com/intel/ipex-llm/blob/main/docs/mddocs/Quickstart/ollama_portable_zip_quickstart.md
IPEX-LLM is archived no more support for newer model like qwen3. What's next?
Thank you, thank you… I couldn't for the life of me get Ollama to work with Intel Arc Pro 50b, I watched your video in less than 5 minutes I had it work, and WOW its fast…. Big Thanks!!
I think, to be fair on the CPU vs GPU comparison, you shoud set "/set parameter num_thread 16" (for a 8 cores cpu?) so it can use 100% of the cpu.
BETTER question: how to set it on Docker?
what if I have an ARC A770 and Linux Mint 22.2?
In comparison to British English American English sounds vulgar.
Like I mentioned in a previous comment, I have a similar spec laptop as yours. I ssked three models a simple proof by Modus Ponens. I got:
Phi4: 14b – 5.09 tokens/s
Phi4-mini: 3.8b – 15.70 tokens/s
Deepseek-r1: 8b – 10.76 tokens/s
So the 11 tokens/s you were getting are pretty much in line with what I'm getting.
I have a Dell laptop with the same CPU, GPU, and NPU. Thanks for the video. There's not a lot of straight forward help for Intel hardware out there.
Would you consider doing the same for the NPU version? Like you I don't think it'll be as fast as the GPU since NP are a prettynew thing, but who knows?
Yeah, that's extremely weak compared to the AMD APU I've been using.
I was genuinely hoping this intel setup would work with me, but that looks like a very weak scenario to me now.
Latest models do not run with IPEX-LLM 2.2 stable , but one of latest beta 2.3 was useable with newest models.
Too bad comfortable response rate is limiting what size model to use, my hardware is allowing even 30b but thats not practical to work with 1.5 tokens /sec
Strange, after you have switched from GPU to CPU-only mode and repeated the test you have a CPU usage of 30% and get around 5.46 tokens/s. If the system would utilise the CPU to 100% capacity, you would get approx. 18 tokens/s. This is 50% faster than reasoning in GPU mode with its 11.83 tokens/s.
Thanks a lot. Can you use a GUI, like AnythingLLM with this?
Muchas gracias, me ha funcionado muy bien con una Intel Arc B570, he ejecutado varios modelos y el resultado es muy satisfactorio, es una tarjeta muy rápida en modelos de 7 y 8B.
Nice video. It's very informative. I wonder why only 12 tokens / s on LLama3.1 8b model. I expected 30+ tokens /s