<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Ubaid Ur Rahman]]></title><description><![CDATA[Ubaid Ur Rahman]]></description><link>https://ubaidurrahman.hashnode.dev</link><image><url>https://cdn.hashnode.com/res/hashnode/image/upload/v1593680282896/kNC7E8IR4.png</url><title>Ubaid Ur Rahman</title><link>https://ubaidurrahman.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Mon, 05 Oct 2026 17:34:46 GMT</lastBuildDate><atom:link href="https://ubaidurrahman.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[ Building a Multi-Tenant WhatsApp Bot for Local Businesses]]></title><description><![CDATA[A restaurant in Karachi doesn't need a custom-built ordering bot. It needs a bot — fast to set up, cheap to run, and easy to hand off to the next restaurant without rewriting code.
That constraint sha]]></description><link>https://ubaidurrahman.hashnode.dev/building-a-multi-tenant-whatsapp-bot-for-local-businesses</link><guid isPermaLink="true">https://ubaidurrahman.hashnode.dev/building-a-multi-tenant-whatsapp-bot-for-local-businesses</guid><dc:creator><![CDATA[Ubaid Ur Rahman]]></dc:creator><pubDate>Sat, 20 Jun 2026 08:07:50 GMT</pubDate><content:encoded><![CDATA[<p>A restaurant in Karachi doesn't need a custom-built ordering bot. It needs <em>a</em> bot — fast to set up, cheap to run, and easy to hand off to the next restaurant without rewriting code.</p>
<p>That constraint shaped the entire architecture of the WhatsApp order bot I built, starting from a single demo (a biryani restaurant) and generalizing it into something any local business could plug into.</p>
<h2>The Wrong Way to Start</h2>
<p>The first working version was hardcoded for one client: one menu, one set of business hours, one phone number, credentials baked into the code. It worked as a demo. It was structurally useless as a product — every new client would mean cloning the repo and editing source code.</p>
<p>That's not a business. That's a one-off project repeated badly.</p>
<h2>The Fix: Config-Driven Architecture</h2>
<p>The rebuild moved every client-specific detail out of code and into a single <code>config.json</code>:</p>
<pre><code class="language-json">{
  "business_name": "Karachi Biryani House",
  "menu": [...],
  "hours": { "open": "11:00", "close": "23:00" },
  "whatsapp_number_id": "...",
  "currency": "PKR",
  "tone": "friendly, casual Urdu-English mix"
}
</code></pre>
<p>The bot logic reads this config at startup and at runtime. Onboarding a new restaurant becomes: write a new config file, point a new WhatsApp Business number at the same deployment. No code changes, no redeploy of core logic.</p>
<p>This is the difference between a project and a product. A project solves one instance of a problem. A product solves the <em>shape</em> of the problem.</p>
<h2>What Was Still Missing</h2>
<p>Getting the config-driven structure right was necessary but not sufficient. A demo bot and a production bot differ in what happens when things go wrong, not in what happens when things go right.</p>
<p><strong>No retry logic.</strong> If a Groq API call failed — rate limit, timeout, network blip — the conversation just died. The customer got no response and no error message. They assumed the bot was broken and left.</p>
<p><strong>No structured logging.</strong> Debugging meant adding <code>print()</code> statements, redeploying, and waiting for the issue to recur. There was no way to see what had happened five minutes ago without it still being on screen.</p>
<p><strong>No state persistence across restarts.</strong> Order-in-progress state lived in memory. A deploy or crash meant every active conversation lost its context mid-order.</p>
<p><strong>Single point of failure on every external call.</strong> Meta's WhatsApp Cloud API, Groq's inference API — any hiccup in either one took down the whole flow with no fallback behavior.</p>
<h2>Fixing It for Production</h2>
<p>The retry pattern is simple but non-negotiable for anything calling an external API in a user-facing flow:</p>
<pre><code class="language-python">def call_with_retry(func, *args, max_retries=3, **kwargs):
    for attempt in range(max_retries):
        try:
            return func(*args, **kwargs)
        except Exception as e:
            if attempt == max_retries - 1:
                raise
            time.sleep(2 ** attempt)
</code></pre>
<p>Exponential backoff, hard cap on attempts, and — critically — a defined behavior on final failure. The bot needs to tell the customer "I'm having trouble right now, try again in a moment" rather than going silent. Silent failure is the worst failure mode in a conversational interface, because the user has no idea whether to wait or walk away.</p>
<p>For state, moving order-in-progress data to Redis instead of in-process memory meant a deploy or restart no longer meant losing every active conversation. This also opened the door to running multiple bot instances behind a load balancer later, since state isn't tied to a single process anymore.</p>
<p>For logging, replacing every <code>print()</code> with structured logging (<code>logging</code> module, timestamped, leveled) meant production issues could be diagnosed from log history instead of needing to reproduce them live.</p>
<h2>What This Pattern Generalizes To</h2>
<p>None of this is specific to WhatsApp bots or restaurants. The same shape applies to any tool meant to serve multiple clients from one codebase:</p>
<ul>
<li><p>Separate configuration from logic completely</p>
</li>
<li><p>Assume every external API call will fail sometimes, and decide in advance what happens when it does</p>
</li>
<li><p>Log enough that you can debug from history, not from live reproduction</p>
</li>
<li><p>Keep state somewhere that survives a restart</p>
</li>
</ul>
<p>The architecture that makes a bot sellable to a second client is rarely about adding features. It's about removing the assumptions that only held true for the first one.</p>
]]></content:encoded></item><item><title><![CDATA[Why Multilingual Embeddings Beat Translation Layers]]></title><description><![CDATA[I ran into the same architectural decision twice while building Islamic AI tools: do you translate the user's query before search, or do you use an embedding model that understands multiple languages ]]></description><link>https://ubaidurrahman.hashnode.dev/why-multilingual-embeddings-beat-translation-layers</link><guid isPermaLink="true">https://ubaidurrahman.hashnode.dev/why-multilingual-embeddings-beat-translation-layers</guid><dc:creator><![CDATA[Ubaid Ur Rahman]]></dc:creator><pubDate>Sat, 20 Jun 2026 08:06:47 GMT</pubDate><content:encoded><![CDATA[<p>I ran into the same architectural decision twice while building Islamic AI tools: do you translate the user's query before search, or do you use an embedding model that understands multiple languages natively?</p>
<p>Both work. They are not the same, and picking the wrong one for your use case costs you either money, latency, or accuracy — sometimes all three.</p>
<h2>The Translation Layer Approach</h2>
<p>This is the obvious first solution. User types in Roman Urdu or Arabic, you call an LLM API to translate it to English, then you embed the English version and search your English-indexed documents.</p>
<pre><code class="language-plaintext">Query → Groq/OpenAI translate → English text → Embed → FAISS search
</code></pre>
<p>It works, and it's fast to implement if your index already exists in English. But it has real costs that don't show up until you're past the demo stage:</p>
<p><strong>Per-query API cost.</strong> Every single search now requires an external API call before you even start searching. At scale, this is not free, and it's not optional — it's on the critical path of every request.</p>
<p><strong>Latency stacking.</strong> Translation typically adds 300ms-1s before search even begins. Users feel this, especially compared to a direct embed-and-search flow.</p>
<p><strong>Translation drift.</strong> LLM translation isn't deterministic. The same Roman Urdu query translated twice can come back slightly different, which means your search results aren't fully reproducible either.</p>
<p><strong>A single point of failure.</strong> If the translation API is down, rate-limited, or slow, your entire search pipeline is down — even though FAISS itself is working fine.</p>
<p>I hit this directly: a translation call that worked for short queries started failing with <code>413 Payload Too Large</code> once I was also stuffing retrieved hadith text into the same API call for answer generation. Two unrelated bottlenecks, same dependency, same outage.</p>
<h2>The Multilingual Embedding Approach</h2>
<p>The alternative: use an embedding model trained across multiple languages from the start, so you never need a translation step at all.</p>
<pre><code class="language-plaintext">Query (any supported language) → Embed directly → FAISS search
</code></pre>
<p>I switched from <code>all-MiniLM-L6-v2</code> (English-only) to <code>paraphrase-multilingual-MiniLM-L12-v2</code> (50+ languages, including Urdu and Arabic). Same model size, same inference speed, no extra API call.</p>
<p>A query like "namaz ka tariqa" gets embedded into roughly the same vector space as its English equivalent "method of prayer" — close enough in cosine similarity for FAISS to retrieve relevant English-language documents directly. No translation hop, no external dependency at query time.</p>
<h2>The Real Trade-off</h2>
<p>The catch is upfront cost, not ongoing cost. Switching embedding models means your entire index is now invalid — every document needs to be re-embedded and the FAISS index rebuilt from scratch. For a corpus of 35,000+ documents, this took about 5-10 minutes on CPU. For a much larger corpus, this could be hours, and it's a cost you pay every time you want to upgrade the embedding model.</p>
<p>There's also an accuracy trade-off worth being honest about: multilingual models are generally weaker than language-specific models at any single language. If 95% of your traffic is English and 5% is Urdu, a dedicated English model with a translation fallback for the 5% might outperform a multilingual model across the board.</p>
<h2>How I'd Decide</h2>
<table>
<thead>
<tr>
<th>Factor</th>
<th>Translation Layer</th>
<th>Multilingual Embeddings</th>
</tr>
</thead>
<tbody><tr>
<td>Query latency</td>
<td>Higher (API call)</td>
<td>Lower (direct embed)</td>
</tr>
<tr>
<td>Cost per query</td>
<td>API cost every search</td>
<td>Free after index built</td>
</tr>
<tr>
<td>Reliability</td>
<td>Depends on external API uptime</td>
<td>Self-contained</td>
</tr>
<tr>
<td>Best for</td>
<td>Mostly single-language, occasional other-language queries</td>
<td>Genuinely mixed-language user base</td>
</tr>
<tr>
<td>Setup cost</td>
<td>Low</td>
<td>Higher (full reindex)</td>
</tr>
</tbody></table>
<p>If you're building for a user base that's <em>genuinely</em> multilingual — not just occasionally typing in another language — multilingual embeddings are the better long-term architecture. The reindex cost is paid once. The translation layer's costs are paid on every single query, forever, for as long as the system runs.</p>
<p>I ended up using both in the same system: multilingual embeddings as the primary search mechanism, with the LLM translation step kept only for generating the final natural-language answer from retrieved context — a task where you genuinely need an LLM in the loop anyway.</p>
]]></content:encoded></item><item><title><![CDATA[Fixing Roman Urdu Search in RAG Pipelines]]></title><description><![CDATA[When I built a semantic search engine over 35,000+ hadiths, the English queries worked well from day one. FAISS retrieved relevant results, the embeddings made sense, the demo looked clean.
Then I tes]]></description><link>https://ubaidurrahman.hashnode.dev/fixing-roman-urdu-search-in-rag-pipelines</link><guid isPermaLink="true">https://ubaidurrahman.hashnode.dev/fixing-roman-urdu-search-in-rag-pipelines</guid><dc:creator><![CDATA[Ubaid Ur Rahman]]></dc:creator><pubDate>Sat, 20 Jun 2026 08:05:15 GMT</pubDate><content:encoded><![CDATA[<p>When I built a semantic search engine over 35,000+ hadiths, the English queries worked well from day one. FAISS retrieved relevant results, the embeddings made sense, the demo looked clean.</p>
<p>Then I tested it the way my actual users would use it: in Roman Urdu.</p>
<p>"aurat ki namaz ka tariqa" returned almost nothing useful. Not because the hadiths weren't there — they were. The embedding model simply didn't understand Roman Urdu as a language. It was treating "namaz" as a meaningless token, not as the Urdu word for prayer.</p>
<p>This is the gap that breaks most RAG systems built for South Asian or Middle Eastern users: the embedding model is trained primarily on English, and Roman Urdu — Urdu written in Latin script — looks like gibberish to it.</p>
<h2>Why This Happens</h2>
<p>A standard embedding model like <code>all-MiniLM-L6-v2</code> is trained on English text. It maps semantically similar English sentences close together in vector space. But "namaz ka tariqa" has no representation in that space — it's not English, and it's not standard Urdu script either. It's a hybrid the model was never trained on.</p>
<p>The naive fix — just embed the Roman Urdu query directly and hope for the best — fails silently. You get results, but they're not the right ones. Worse, there's no error to catch. The system looks like it's working.</p>
<h2>The Fix: Translation Before Embedding</h2>
<p>The pipeline needs an explicit step before the query ever touches FAISS:</p>
<pre><code class="language-plaintext">User query → Detect language → If Roman Urdu: translate to English → Embed → Search
</code></pre>
<p>I used Groq's API with a fast model to handle translation. The key detail most people get wrong here: <strong>detecting Roman Urdu is not the same as detecting non-English text.</strong></p>
<p>My first attempt used a simple ASCII-ratio check — if 85% of characters were ASCII, treat it as English. This failed immediately, because Roman Urdu <em>is</em> ASCII. "aurat ka namaz ka tariqa" passes an ASCII check with flying colors while being completely unintelligible to an English-trained model.</p>
<p>The actual fix was a marker-word approach:</p>
<pre><code class="language-python">ROMAN_URDU_MARKERS = {
    "ka", "ki", "ke", "ko", "ne", "se", "mein", "par", "hai",
    "hain", "namaz", "zakat", "roza", "hadith", "sunnat",
    "wuzu", "salat", "dua", "masjid", "quran", "allah"
}

def is_likely_english(query: str) -&gt; bool:
    tokens = set(query.lower().split())
    if tokens &amp; ROMAN_URDU_MARKERS:
        return False
    ascii_ratio = sum(c.isascii() for c in query) / len(query)
    return ascii_ratio &gt;= 0.85
</code></pre>
<p>If the query contains common Roman Urdu function words or domain vocabulary, it routes to translation regardless of script. Only then do you fall back to the ASCII heuristic for genuinely ambiguous cases.</p>
<h2>The Second Layer: Multilingual Embeddings</h2>
<p>Translation solves the query side, but it adds latency and an external API dependency for every single search. The better long-term fix is switching the embedding model itself.</p>
<p>I moved from <code>all-MiniLM-L6-v2</code> to <code>paraphrase-multilingual-MiniLM-L12-v2</code> — same architecture family, same inference speed, but trained across 50+ languages including Urdu and Arabic. This means a Roman Urdu or Arabic-script query can match an English-embedded document directly, without a translation hop at all.</p>
<p>The trade-off: this requires rebuilding your FAISS index from scratch, since the vector space itself changes with the model. For a 35,000-document corpus, that's a one-time cost of about 5-10 minutes on CPU — trivial compared to the alternative of running translation calls on every query forever.</p>
<h2>What I'd Tell Someone Starting This</h2>
<p>If you're building a RAG system for a non-English-first user base, don't treat multilingual support as a feature to add later. Decide upfront whether you're translating queries or using multilingual embeddings — they solve the same problem differently, and retrofitting either one means touching your entire ingestion and query pipeline.</p>
<p>And test with real queries from real users, in the script and phrasing they'd actually type. Clean English test queries will tell you nothing about whether your system works for the people you're actually building it for</p>
]]></content:encoded></item></channel></rss>