# ai.rs — Custom AI Company in Serbia, Europe > Custom AI solutions, LLM fine-tuning, and AI assistant development for e-commerce and retail businesses. Based in Belgrade, Serbia. Serving clients across Europe and worldwide. ## What We Do ai.rs is a custom AI company that builds domain-specific AI assistants for businesses. We fine-tune open-source large language models (Qwen, Llama, Mistral) on your product data, deploy them on dedicated GPU servers, and connect them via RAG pipelines for real-time pricing and availability. We are an LLM fine-tuning service specializing in e-commerce AI, product recommendation engines, and AI-powered sales assistants. ## Location - Based in Belgrade, Serbia, Europe - .rs is the country-code top-level domain for Serbia - Serving clients in Europe, the Balkans, and worldwide - Remote-first — we work with businesses everywhere ## Services - Custom AI sales assistants trained on your product catalog - LLM fine-tuning service (Qwen3-8B, Llama 3, Mistral 7B, and other open-source models) - RAG pipeline development for real-time product data - Safety and guardrail training (~26,000 training examples, 94% attack resistance) - Managed AI hosting on dedicated GPU infrastructure - AI consulting for e-commerce and retail ## Models We Work With - Qwen3-8B, Qwen 2.5 (Alibaba) — excellent multilingual support - Llama 3.1 8B, Llama 3 70B (Meta) — strong English performance - Mistral 7B, Mixtral (Mistral AI) — efficient inference - Any open-source model on Hugging Face — we evaluate and recommend the best fit ## Pricing Three tiers: Starter (up to 200 products), Business (up to 1,000 products), Enterprise (unlimited). Performance Partnership (revenue-share) also available. See [How It Works](https://ai.rs/how-it-works#pricing) for current pricing. ## Follow Us - [Telegram Channel](https://t.me/www_ai_rs): AI news, articles, and practical insights ## Key Pages - [How It Works](https://ai.rs/how-it-works): Full process, pricing, FAQ - [AI Readiness Assessment](https://ai.rs/ai-readiness): Free 2-minute business quiz - [All Articles](https://ai.rs/filter): Browse and filter all articles by category and tag - [Search Articles](https://ai.rs/filter/search/rag): Fulltext search across all articles (replace "rag" with your query) - [Contact](https://ai.rs/contact): Get started - [Full site content for LLMs](https://ai.rs/llms-full.txt): Complete documentation ## API — Article Search - Search: `https://ai.rs/api/articles.php?q={query}` — fulltext search with relevance ranking - Filter by category: `https://ai.rs/api/articles.php?category=fundamentals` - Combine: `https://ai.rs/api/articles.php?q=rag&category=fundamentals` - Returns JSON array of articles sorted by relevance ## Articles — Business - [Why Your Business Needs Its Own AI Model](https://ai.rs/ai-for-business/why-your-business-needs-its-own-ai-model): Why generic AI like ChatGPT isn't enough for your business and how a custom AI model trained on your products becomes an unfair competitive advantage. - [Your AI Sales Associate That Never Sleeps](https://ai.rs/ai-for-business/your-ai-sales-associate-that-never-sleeps): How an AI sales assistant works 24/7, speaks your customer's language, handles 100 customers at once, and costs a fraction of a single employee. - [How Custom AI Increases Sales Conversions](https://ai.rs/ai-for-business/how-custom-ai-increases-sales-conversions): How AI sales assistants increase conversion rates by 10-30% and average order value by 10-25% through guided selling, cross-selling, and 24/7 availability. - [What is RAG: Why Your AI Always Has the Right Price](https://ai.rs/ai-for-business/what-is-rag-why-your-ai-always-has-the-right-price): How RAG technology ensures your AI assistant always quotes the right price, knows about new products instantly, and never recommends discontinued items. - [ChatGPT vs Your Own Model: What's the Difference?](https://ai.rs/ai-for-business/chatgpt-vs-your-own-model-whats-the-difference): Side-by-side comparison of ChatGPT API vs your own AI model covering data privacy, cost, brand voice, product knowledge, and competitor control. - [What Does It Cost? The Real Numbers Behind Custom AI](https://ai.rs/ai-for-business/what-does-custom-ai-cost-real-numbers): Complete cost breakdown of custom AI for business: hardware, training, monthly operating costs, break-even timeline, and comparison with API and human alternatives. - [From Zero to Live: What Getting Custom AI Actually Looks Like](https://ai.rs/ai-for-business/from-zero-to-live-getting-custom-ai): Week-by-week timeline of implementing custom AI for your business: what you provide, what happens, and how you go from zero to a live AI assistant in 4 weeks. - [5 Things Your AI Should Never Do: Safety for Business Owners](https://ai.rs/ai-for-business/five-things-your-ai-should-never-do): The 5 critical safety rules every business AI needs: never invent products, change prices, go off-topic, reveal instructions, or recommend competitors. - [Will AI Replace My Sales Team? (No — Here's Why)](https://ai.rs/ai-for-business/will-ai-replace-my-sales-team): Why AI augments your sales team instead of replacing them. What AI does better, what humans always will, and how the combination multiplies your sales capacity. - [RGB Delight: Raspberry Pi 2 + Arduino Nano + WS2812b Ambilight with Hyperion](https://ai.rs/ai-developer/rgb-delight-raspberry-pi-arduino-hyperion-ambilight): Build a DIY Ambilight TV clone with Raspberry Pi 2, Arduino Nano, WS2812b RGB LED strip, and Hyperion on OpenElec. Includes Arduino Adalight sketch, wiring, and troubleshooting. - [SEO Is Dead. Your Rankings Don't Matter Anymore.](https://ai.rs/ai-for-business/seo-is-dead-rankings-dont-matter): LinkedIn lost 60% of B2B traffic while rankings held steady. AI search is killing clicks. Here is what businesses need to do differently. - [Building an Email List That Survives the Algorithm](https://ai.rs/ai-for-business/building-email-list-that-survives-algorithm): Email is the only audience channel no platform can take away. Practical guide to building a B2B email list: lead magnets, send cadence, metrics, and tech stack. - [You're Sitting on a Goldmine of AI Training Data](https://ai.rs/ai-for-business/youre-sitting-on-a-goldmine-of-ai-training-data): Your business already has the training data for a custom AI model. Learn how to turn chatbot logs, call recordings, product catalogs, and support tickets into a production-ready dataset. - [Your Competitors Aren't Using AI Yet — Make That Your Advantage](https://ai.rs/ai-for-business/competitors-arent-using-ai-yet-your-advantage): New research shows 94% of business tasks could be handled by AI, but only 33% actually are. Here's how smart business owners are using that gap to win. - [AI Won't Replace Your Team — But a Team Using AI Will Replace Yours](https://ai.rs/ai-for-business/ai-wont-replace-your-team): 57% of workplace AI use is augmentation, not replacement. Learn how to make your existing team 2-5x more productive with AI — a practical 90-day playbook. - [100% ROI in 24 Hours: Nvidia B200 Replaced a $35,000 AI API Bill in a Single Day](https://ai.rs/ai-for-business/100-percent-roi-in-24-hours-nvidia-b200-replaced-35000-ai-api-bill): How we cut AI text generation costs from $35,000 to $180 by self-hosting Qwen3.5 on an Nvidia B200 GPU. A 194x cost reduction case study for batch AI processing at scale. - [100% Human Key-Pressed: Share This Email Signature](https://ai.rs/ai-for-business/100-percent-human-key-pressed-email-signature): Share this three-line email signature: 100% human key-pressed, 0% machine-generated, 99% naturally imperfect. Plus why Gmail Smart Compose and Outlook Copilot are flattening your voice. ## Articles — Beginner - [What Is AI, Really? Machine Learning, Deep Learning, and Generative AI Explained](https://ai.rs/learn-ai/what-is-ai-really): Clear, jargon-free explanation of AI, machine learning, deep learning, and generative AI — what each term means, how they relate, and why it matters. - [How ChatGPT Actually Works (In Plain English)](https://ai.rs/learn-ai/how-chatgpt-actually-works): A plain-English explanation of how ChatGPT works — from training on the internet to predicting one word at a time — without any technical jargon. - [AI Prompting 101: How to Get Better Answers Every Time](https://ai.rs/learn-ai/ai-prompting-101): Learn how to write better AI prompts with five practical rules, common mistakes to avoid, and power techniques that dramatically improve your results. - [What Are AI Hallucinations and How to Spot Them](https://ai.rs/learn-ai/what-are-ai-hallucinations): What AI hallucinations are, why they happen, how to spot them, and practical strategies to reduce the risk of acting on made-up information. - [Local vs Cloud AI: What's the Difference?](https://ai.rs/learn-ai/local-vs-cloud-ai): A clear comparison of local and cloud AI — what each means, how they differ on privacy, cost, and quality, and when to use which approach. - [How to Pick the Right AI Tool for You](https://ai.rs/learn-ai/how-to-pick-the-right-ai-tool): Honest comparison of ChatGPT, Claude, Gemini, Copilot, and Perplexity — how to pick the right AI tool based on what you'll actually use it for. - [What Is Fine-Tuning? Teaching AI New Tricks](https://ai.rs/learn-ai/what-is-fine-tuning): A beginner-friendly explanation of AI fine-tuning — what it is, how it works, what it costs, and when it makes sense for your business or project. - [AI Privacy and Safety: What Every User Should Know](https://ai.rs/learn-ai/ai-privacy-and-safety-basics): Practical guide to AI privacy and safety — where your data goes, what not to share, how bias works, and concrete steps to use AI tools responsibly. ## Articles — Developer - [What is an LLM and How to Deploy It on Your Website](https://ai.rs/ai-developer/what-is-llm-how-to-deploy): Learn what a large language model is and how to deploy one on your business website with practical steps, hardware requirements, and performance benchmarks. - [Why Train Your Own LLM: Advantages for Business](https://ai.rs/ai-developer/why-train-your-own-llm): Discover why fine-tuning your own LLM with LoRA gives your business a domain expert AI assistant at a fraction of the cost of API-based solutions. - [How LLM Can Transform Sales and Customer Support](https://ai.rs/ai-developer/how-llm-transforms-sales-support): Explore how a self-hosted LLM transforms e-commerce sales and customer support with instant responses, personalized recommendations, and 24/7 availability. - [What is RAG and Why Your AI Needs It](https://ai.rs/ai-developer/what-is-rag-why-your-ai-needs-it): Learn how Retrieval-Augmented Generation (RAG) gives your AI assistant accurate, real-time product knowledge without retraining the model. - [vLLM vs Ollama: When Do Advanced Serving Frameworks Win?](https://ai.rs/ai-developer/vllm-vs-ollama-serving-frameworks): Head-to-head benchmarks comparing Ollama and vLLM for LLM inference. Single-user vs multi-user performance, setup complexity, and when to switch. - [BM25 vs Embeddings: Why Keyword Search Still Wins for Product Data](https://ai.rs/ai-developer/bm25-vs-embeddings-keyword-search): Production benchmarks comparing BM25 keyword search vs neural embeddings for e-commerce product retrieval. BM25 wins 92% vs 78% on exact matches. - [Quantization Methods Compared: GGUF, AWQ, GPTQ, EXL2, NVFP4](https://ai.rs/ai-developer/quantization-methods-compared): Comprehensive comparison of GGUF, GPTQ, AWQ, EXL2, and NVFP4 quantization formats with benchmarks for speed, quality, and non-English language support. - [Building a Domain-Specific AI Product Recommender](https://ai.rs/ai-developer/building-domain-specific-product-recommender): Step-by-step guide to building an AI-powered product recommender using fine-tuning and RAG that provides contextual, expert-level recommendations. - [LLM Security: From 17% to 94% Attack Resistance](https://ai.rs/ai-developer/llm-security-attack-resistance): How targeted safety training with just 275 samples improved an LLM's attack resistance from 17% to 94%, covering prompt injection, jailbreaks, and data manipulation. - [The GPU Memory Wall: Why Inference Hardware Matters](https://ai.rs/ai-developer/gpu-memory-wall-inference-hardware): Deep dive into why GPU inference is memory-bound with 98% idle cores, how quantization helps, and why purpose-built ASICs are 10-100x faster. - [When the Memory Wall Disappears: What Actually Bottlenecks LLM Inference on Modern GPUs](https://ai.rs/ai-developer/memory-wall-disappears-llm-inference-bottlenecks): We pinned a quantized 135M model in GPU L2 cache to eliminate the memory wall. What replaced it — kernel dispatch overhead — explains why inference ASICs exist. - [From Edge AI to Custom LLMs: How On-Device Intelligence Evolved](https://ai.rs/ai-developer/edge-ai-kendryte-k210-to-custom-llms): From the $20 Kendryte K210 edge AI camera to fine-tuned 8B parameter LLMs — how on-device intelligence evolved from object detection to conversational AI assistants. - [Sunday Project: Ambilight Using Raspberry Pi and RGB LED Strip](https://ai.rs/ai-developer/ambilight-raspberry-pi-rgb-led-strip): Build a DIY Ambilight TV backlight with Raspberry Pi 2, Arduino Nano, and WS2812b RGB LED strip using Hyperion on XBMC/Kodi. - [How to Flash Hackaday Badge 2018](https://ai.rs/ai-developer/how-to-flash-hackaday-badge-2018): Step-by-step guide to flashing the Hackaday Belgrade 2018 conference badge firmware using PICkit3 and MPLAB X IDE. - [How to Block Ads Using Pi-Hole and NanoPi Neo2](https://ai.rs/ai-developer/pi-hole-nanopi-neo2-ad-blocking): Set up network-wide ad blocking with Pi-Hole on a NanoPi Neo2 ARM computer with Gigabit Ethernet for fast, silent DNS-level ad filtering. - [Mercury 2: The First Reasoning Diffusion LLM — 1,000 Tokens/sec](https://ai.rs/ai-developer/mercury-2-diffusion-reasoning-llm): Mercury 2 from Inception Labs is the first reasoning diffusion LLM, generating 1,000 tokens/sec by producing tokens in parallel. Here's how it works and what it means for developers. - [Claude Code Remote Control: Continue Coding Sessions from Your Phone](https://ai.rs/ai-developer/claude-code-remote-control-mobile): Claude Code Remote Control lets you continue local coding sessions from your phone or browser. Here's how to set it up, real-world use cases, and how it compares to cloud-based coding. - [How to Implement llms.txt — The Developer's Guide](https://ai.rs/ai-developer/how-to-implement-llms-txt): A practical guide to implementing llms.txt — the Markdown file that helps AI systems understand your website. Format, examples, and honest assessment of who reads it. - [Llama 4 vs Qwen 3.5 vs Gemma 3: Which Open Model Should You Deploy?](https://ai.rs/ai-developer/llama-4-vs-qwen-3-5-vs-gemma-3-compared): Head-to-head benchmarks comparing Llama 4 Scout, Qwen 3.5, and Gemma 3 on reasoning, coding, multilingual, inference speed, and VRAM requirements for self-hosted deployment. - [Will This LLM Fit My GPU? VRAM Requirements for Every Model Size](https://ai.rs/ai-developer/will-llm-fit-my-gpu-vram-requirements): Check if an LLM fits your GPU before downloading. VRAM formula, model size tables for 8-32 GB GPUs, and a one-command tool to check any Hugging Face model. - [LLM Post-Training Explained: SFT, DPO, and GRPO](https://ai.rs/ai-developer/llm-post-training-explained): Understand the three stages of LLM post-training: Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Group Relative Policy Optimization (GRPO). Practical guide with pros, cons, and tools. - [Synthetic Data for Fine-Tuning: How to Generate Your Own Training Set](https://ai.rs/ai-developer/synthetic-data-for-fine-tuning): Learn how to generate thousands of high-quality training samples for LLM fine-tuning using synthetic data pipelines. Covers seed prompts, LLM-as-judge, filtering, and practical tools. - [Gemma 4 vs Qwen 3.5 vs Llama 4: Updated Benchmarks, New Leader](https://ai.rs/ai-developer/gemma-4-vs-qwen-3-5-vs-llama-4-compared): Gemma 4 benchmarks obliterate Gemma 3: 89% on AIME math, 80% on LiveCodeBench, 84% on GPQA. The MoE variant matches 31B quality with 4B active params. Apache 2.0 licensed. - [Gemma 4 LoRA Fine-Tuning on RTX 5090: What Works and What Doesn't](https://ai.rs/ai-developer/gemma-4-lora-fine-tuning-rtx-5090): Gemma 4's MoE 26B-A4B uses 3D fused expert tensors that bitsandbytes can't quantize yet — blocking QLoRA on 32 GB GPUs. Dense Gemma 4 models work fine. We explain the blocker, workarounds, and when a fix is coming. - [Why Every AI Engineer Should Learn Classical Chinese](https://ai.rs/ai-developer/classical-chinese-agent-memory-compression): Benchmarking Classical Chinese (Wenjian) vs AAAK vs English as an agent-memory format. 24% token savings at 96% retrieval — and the surprising lesson about which model to evaluate on. - [Qwen 3.6 27B: a Local Coding Model You Can Actually Run](https://ai.rs/ai-developer/qwen-3-6-27b-local-coding-model): Qwen 3.6 27B is the first open coding model that runs on a single 24GB GPU and gets within 4 points of Claude Opus 4.6 on SWE-bench. Here's how to run it. - [How to Run Qwen3-Coder 30B-A3B on RTX 5090 with Ollama](https://ai.rs/ai-developer/qwen3-coder-30b-a3b-rtx-5090-ollama): Step-by-step guide to running Qwen3-Coder 30B-A3B locally on RTX 5090 with Ollama. 231 TPS via MoE, 64K context with q8_0 KV cache, and the Modelfile gotchas (RENDERER, PARSER) you won't find in the docs. - [The KV Cache: the trick that makes LLMs fast — and slow](https://ai.rs/ai-developer/kv-cache-makes-llms-fast-and-slow): Why the KV cache makes LLM generation fast but long context slow: the memory it stores and the bandwidth it streams both grow with every token, until you hit the wall. - [Prompt Processing vs Token Generation: the Two Speeds of an LLM](https://ai.rs/ai-developer/prompt-processing-vs-token-generation): Prompt Processing (PP/prefill) vs Token Generation (TG/decode): the two speeds of an LLM, why they differ (compute- vs bandwidth-bound), and how to read benchmark numbers correctly. - [AI Workstation Comparison: RTX 5090 vs GB10 (HP ZGX)](https://ai.rs/ai-developer/rtx-5090-vs-gb10-hp-zgx): RTX 5090 vs GB10 (HP ZGX) for local AI: 32 GB/1.8 TB/s vs 128 GB/273 GB/s. Memory, inference speed (PP vs TG), MoE, long context, and power consumption compared — and which to buy. - [Qwen-AgentWorld: the Open Language World Model for AI Agents](https://ai.rs/ai-developer/qwen-agentworld-language-world-model): Qwen-AgentWorld is an open-weight language world model that simulates agent environments across seven domains (MCP, Search, Terminal, SWE, Android, Web, OS) for sim-RL and agent training. The 397B MoE tops AgentWorldBench, edging GPT-5.4. - [Mixture of Experts (MoE), Explained](https://ai.rs/ai-developer/mixture-of-experts-explained): Mixture of Experts (MoE) explained: dense vs MoE, total vs active parameters, why it gives a giant model's quality at a small model's speed, and the memory trade-off. - [4-Bit Quantization Decoded: INT4 QAT, MXFP4, and NVFP4](https://ai.rs/ai-developer/int4-qat-mxfp4-nvfp4-quantization): INT4 QAT vs MXFP4 vs NVFP4 explained: 4-bit floating point (E2M1), block sizes and scales, why NVFP4 is more accurate, and which 4-bit format to use on Blackwell GPUs. - [What a 256K (or 1M) Context Window Actually Costs You](https://ai.rs/ai-developer/context-window-cost): What a 256K or 1M context window actually costs: KV cache memory (grows linearly), prompt-processing compute, and per-token bandwidth — plus how to cut it with MLA, KV quantization, prompt caching, and RAG. - [DeepSeek-V4-Flash on Two GB10s: 304B Params, 1M Context, 83 Watts](https://ai.rs/ai-developer/deepseek-v4-flash-304b-two-gb10): DeepSeek-V4-Flash (304B MoE, 1M context) on two GB10 / DGX Spark boxes: 88 tok/s decode on 83 W, measured — a self-hosted coding agent that isn't a toy. - [Self-Hosted Claude Code: What Each Memory Tier Actually Buys You](https://ai.rs/ai-developer/self-hosted-claude-code-memory-tiers): Self-hosting a coding agent: the tool is the easy half. What 32 GB, 128 GB, and 256 GB each buy you in model capability, with measured throughput and power. ## Articles — News - [Mercury 2: Hands-On With the World's Fastest Reasoning LLM](https://ai.rs/ai-for-business/mercury-2-hands-on-fastest-reasoning-llm): Hands-on testing of Mercury 2, the first commercial diffusion LLM. Speed benchmarks, tool use, streaming, structured output, and the max_tokens quirk you need to know. - [Qwen 3.5: 35B Knowledge at 4B Speed — Better Than GPT-5?](https://ai.rs/ai-for-business/qwen-3-5-35b-knowledge-4b-speed-better-than-gpt-5): Qwen 3.5 ships 8 models from 0.8B to 397B params with Mixture of Experts architecture. We cover exact VRAM requirements, MoE trade-offs, benchmarks, and how it compares to GLM-5, DeepSeek, and Kimi. - [Claude Mythos Preview: Why Anthropic Locked Its Best Security Model Behind a Wall](https://ai.rs/ai-for-business/claude-mythos-glasswing-why-gated): Claude Mythos Preview found a 27-year-old OpenBSD vulnerability and beats Opus 4.6 on CyberGym 83% to 67%. We break down Project Glasswing access, the 12 founding partners, the pricing, and why Anthropic isn't selling it to you. - [Meta Unveils Muse Spark: First Model From Superintelligence Labs](https://ai.rs/ai-for-business/meta-muse-spark-msl-multimodal-reasoning): Meta Superintelligence Labs launches Muse Spark — a multimodal reasoning model with visual chain-of-thought, tool-use, and parallel multi-agent Contemplating mode. Live on meta.ai today. - [[Restored] Claude Fable 5 Is Free on Your Subscription Until June 22 — Here's How to Switch It On](https://ai.rs/ai-for-business/claude-fable-5-free-test-window-how-to-switch): Claude Fable 5 is free on Pro, Max, Team, and Enterprise plans through June 22, 2026. Here's how to switch to Fable 5 in the app, mobile, and Claude Code — and what happens after June 23. - [Kimi K2.6 Explained: a Trillion-Parameter Open Model](https://ai.rs/ai-for-business/kimi-k2-6-explained): Kimi K2.6 is Moonshot AI's 1T-parameter open-weight MoE (32B active) that edges GPT-5.4 on SWE-Bench Pro. Architecture, memory/VRAM, deployment, and benchmarks explained. - [Apricot Jam: Fable 5 vs Sonnet 5 — Which AI Makes the Better Retro Game?](https://ai.rs/ai-for-business/fable-5-vs-sonnet-5-retro-game-bakeoff): Fable 5 vs Sonnet 5: two Claude models one-shot a playable remake of the 1991 arcade game Apricots in 2D and 3D. Play all four builds, then see who won. - [Claude Opus 5: What's New](https://ai.rs/ai-for-business/claude-opus-5-whats-new): Claude Opus 5 (July 24, 2026): near-Fable 5 intelligence at half the price. Benchmarks, self-verification gains, capped cyber capability, pricing, and whether to switch.