Run Ollama on Your Intel Arc GPU

Ready, click the button in the top right corner to generate summary
AI thinking...



You want to run Ollama but you have an Intel Arc graphics card? I have some great news for you. Introducing IPEX-LLM for all your inference needs. You will be chatting with silicon in no time. This video will walk you how to use the Ollama Portable Zip on Intel GPU with IPEX-LLM. Enjoy!

0:00 Intro (Intel Arc Fans Rejoice)
0:55 Getting Started with IPEX-LLM
3:07 DEMO – Run Ollama on Intel Arc
5:43 DEMO – Initial Tests
11:03 DEMO – Testing Immediate Command Lists
12:08 DEMO – Compare GPU and CPU Performance

Run Ollama Portable Zip on Intel GPU with IPEX-LLM
https://github.com/intel/ipex-llm/blob/main/docs/mddocs/Quickstart/ollama_portable_zip_quickstart.md

Previous Article Code gives hints! Affordable Android phone can support Apple AirDrop file transfer across platforms

14 Comments

  1. Thank you, thank you… I couldn't for the life of me get Ollama to work with Intel Arc Pro 50b, I watched your video in less than 5 minutes I had it work, and WOW its fast…. Big Thanks!!

  2. I think, to be fair on the CPU vs GPU comparison, you shoud set "/set parameter num_thread 16" (for a 8 cores cpu?) so it can use 100% of the cpu.

  3. Like I mentioned in a previous comment, I have a similar spec laptop as yours. I ssked three models a simple proof by Modus Ponens. I got:

    Phi4: 14b – 5.09 tokens/s
    Phi4-mini: 3.8b – 15.70 tokens/s
    Deepseek-r1: 8b – 10.76 tokens/s

    So the 11 tokens/s you were getting are pretty much in line with what I'm getting.

  4. I have a Dell laptop with the same CPU, GPU, and NPU. Thanks for the video. There's not a lot of straight forward help for Intel hardware out there.

    Would you consider doing the same for the NPU version? Like you I don't think it'll be as fast as the GPU since NP are a prettynew thing, but who knows?

  5. Yeah, that's extremely weak compared to the AMD APU I've been using.

    I was genuinely hoping this intel setup would work with me, but that looks like a very weak scenario to me now.

  6. Latest models do not run with IPEX-LLM 2.2 stable , but one of latest beta 2.3 was useable with newest models.

    Too bad comfortable response rate is limiting what size model to use, my hardware is allowing even 30b but thats not practical to work with 1.5 tokens /sec

  7. Strange, after you have switched from GPU to CPU-only mode and repeated the test you have a CPU usage of 30% and get around 5.46 tokens/s. If the system would utilise the CPU to 100% capacity, you would get approx. 18 tokens/s. This is 50% faster than reasoning in GPU mode with its 11.83 tokens/s.

  8. Muchas gracias, me ha funcionado muy bien con una Intel Arc B570, he ejecutado varios modelos y el resultado es muy satisfactorio, es una tarjeta muy rápida en modelos de 7 y 8B.

  9. Nice video. It's very informative. I wonder why only 12 tokens / s on LLama3.1 8b model. I expected 30+ tokens /s