business

Verdict

Submitted 7/26/2026, 1:09:12 PM · Completed 7/26/2026, 1:11:11 PM

6.8
pivot
The idea

Ask HN: Should I Combine Global Knowledge, Internet Search, and User RAG

Show original source text →
I'm building a SaaS platform in Sri Lanka that handles documents and other sensitive data. Each user can upload their own documents and information, and the platform uses RAG to answer questions based on that user's data. That part makes sense to me. My main concern is what happens when the user hasn't uploaded enough information. I still want the LLM to provide accurate answers using reliable information from the internet (or from a curated knowledge base), with proper citations. These are the two architectures I'm considering: Option 1: Base LLM (OpenAI/Anthropic via Azure AI Foundry or Amazon Bedrock) ↓ Platform RAG (global knowledge base managed by us) ↓ User-specific RAG In this approach, we maintain a global knowledge base that we (the platform admins) curate and update. Every user can access this shared knowledge, while their own uploaded documents are searched through their personal RAG. Option 2: Open-source LLM ↓ Fine-tuned on Sri Lankan/domain-specific data ↓ User-specific RAG Here, we fine-tune an open-source model using Sri Lankan or domain-specific data, and each user still has their own RAG for their private documents. My concerns are: Is fine-tuning actually the right solution here, or is it unnecessary? Is a global/shared RAG a better approach than fine-tuning? How would you design this architecture if you wanted: Accurate answers from domain knowledge User-private document search Citations/sources Good scalability for thousands of users I'm leaning toward Option 1 because fine-tuning seems expensive, time-consuming, and I have no experience with it yet. However, I'm not sure if I'm thinking about this correctly. I'd really appreciate hearing how others would approach this problem
TRIZ inventive level: 3/5· Principles: parameter changes, mechanical interaction
Synthesis verdict
**Pivot**: The SaaS platform idea has potential, but it requires adjustments to mitigate significant risks and enhance its value proposition. The hybrid RAG architecture (Option 1) is the most feasible approach, offering a balance of accuracy, scalability, and cost-efficiency. However, regulatory compliance and dependence on external LLMs pose high-impact risks that need to be addressed.

Strengths

  • Clear value proposition for secure document management and AI-driven insights in Sri Lanka's growing digital economy
  • Hybrid RAG architecture (Option 1) leverages existing LLMs, reducing costs and complexity
  • Scalable SaaS model with potential for healthy gross margins (60-70%)
  • Defensible niche due to integrated RAG and potential for exclusive data partnerships

Weaknesses

  • Dependence on external LLMs poses risks related to pricing changes, service interruptions, or policy shifts
  • Regulatory compliance in Sri Lanka is uncertain and potentially costly
  • User trust in data privacy and accuracy is crucial and challenging to establish

Best angle

Refine the platform to focus on compliance and data privacy, exploring local regulations and implementing robust security measures to mitigate risks and build user trust.

Panel verdicts

Viability

meta/llama-4-maverick-17b-128e-instruct

7.0

A solo or 2-person team can build a basic version of the SaaS platform within 4-12 weeks, but it will likely require significant compromises on features or quality.

Building a SaaS platform with RAG capabilities for document handling and sensitive data is feasible for a solo or 2-person team within 4-12 weeks, but with significant complexity. The main challenge lies in implementing the RAG architecture, particularly the global knowledge base and user-specific RAG. Option 1, using a base LLM with a global knowledge base and user-specific RAG, is more feasible as it leverages established LLMs like OpenAI/Anthropic via Azure AI Foundry or Amazon Bedrock. However, curating and updating the global knowledge base will require significant effort. Option 2, fine-tuning an open-source LLM, is more complex and time-consuming, especially for a team without prior experience in fine-tuning LLMs. The team will need to weigh the trade-offs between the two options, considering factors like development time, cost, and the quality of results. For scalability, using a cloud-based infrastructure and designing a robust architecture will be crucial. Citations/sources can be implemented by tracking the sources used by the LLM, but this will add to the complexity.

Competition

nvidia/nemotron-3-nano-omni-30b-a3b-reasoning

7.0

A locally curated, continuously updated global knowledge base combined with per‑user RAG offers a clearer, more durable moat than fine‑tuning, especially for regulated or region‑specific content.

The core need - delivering accurate answers from a mix of proprietary documents and reliable external or curated knowledge - is already addressed by several established platforms. Microsoft 365 Copilot, Notion AI, and enterprise search solutions such as Glean all combine a large language model with a searchable index of an organization's files plus external web or curated data, and they provide citation features. In the Sri Lankan market, local document‑management services (e.g., government e‑services portals) exist, but they rarely embed generative AI with robust citation. Your Option 1 leverages a managed base LLM (OpenAI/Anthropic via Azure or Bedrock) and layers a globally curated knowledge base that you control, plus per‑user RAG for private uploads. This mirrors the architecture of existing AI‑augmented productivity suites, so the differentiation rests on the depth and relevance of the curated knowledge base (e.g., Sri Lankan laws, cultural context, regulated industry data) and on strict data‑privacy compliance for local users. Fine‑tuning an open‑source model (Option 2) would indeed be costly, time‑consuming, and offers limited advantage over a well‑maintained RAG pipeline, especially given the rapid evolution of base models. Scalability is achievable with Option 1 because the heavy lifting (model inference) is handled by the cloud provider, while your team only needs to maintain the knowledge base and user‑specific indexes. The durability of the moat depends on your ability to keep the curated data current, enforce citation standards, and differentiate on local relevance rather than raw model performance. If you can secure exclusive data partnerships or government‑backed content, the differentiation becomes defensible; otherwise, you compete on price and feature parity with the aforementioned platforms.

Risk

nvidia/llama-3.3-nemotron-super-49b-v1

3.0

Dependence on external LLMs and untested regulatory compliance in Sri Lanka pose immediate, high-impact risks.

The proposed SaaS platform faces significant risks, primarily due to its handling of sensitive data and reliance on external AI services. **Regulation** is a major killer: Sri Lanka's data protection laws, though evolving, may impose stringent requirements on storing and processing sensitive user data, potentially necessitating costly compliance measures. **Platform Risk** is another critical factor; dependence on external LLMs (OpenAI/Anthropic via Azure or Amazon) means the platform is at the mercy of these providers' pricing changes, service interruptions, or policy shifts, which could abruptly render the service unviable or force a costly migration. **Churn** due to **No-Budget Customers** isn't as immediate but could plague sustainability if the target market lacks willingness to pay for the added value of curated global knowledge access alongside personal document search.

Monetization

mistralai/mistral-nemotron(fallback #1)

7.0

Option 1 offers a more scalable and cost-effective solution, making it the better choice for immediate implementation and monetization.

The idea has a clear value proposition for users needing secure document management and AI-driven insights. Option 1 is more practical for immediate scalability and cost-efficiency, leveraging existing LLM providers and a curated global knowledge base. This reduces upfront costs and complexity compared to fine-tuning an open-source model. The global RAG ensures accurate, cited answers from reliable sources, while user-specific RAG handles private documents. Monetization can follow a tiered SaaS model: a free tier with basic features, a mid-tier for advanced RAG capabilities, and an enterprise tier for custom integrations and priority support. Pricing could range from $10/user/month for the mid-tier to $50/user/month for enterprise. Conversion paths include a freemium model with clear upsell opportunities. Gross margins should be healthy (60-70%) given the scalable nature of SaaS. Key concerns are data privacy compliance and ensuring the global knowledge base remains up-to-date and accurate.

Market

mistralai/mistral-small-4-119b-2603(fallback #2)

7.0

A hybrid RAG architecture with a curated global knowledge base and user-specific retrieval is the optimal balance of accuracy, scalability, and cost for Sri Lanka's document-focused SaaS market.

The SaaS platform addresses a clear unmet need in Sri Lanka's growing digital economy: secure, localized document management with AI-driven insights. The target audience includes SMEs, legal/finance professionals, and government agencies handling sensitive data - sectors with documented demand for compliance and efficiency tools. Sri Lanka's IT services market is valued at ~$1B, with 10-15% annual growth, and businesses increasingly adopt cloud solutions post-pandemic. However, the market size for niche RAG-based document platforms is smaller: ~5,000 potential enterprise users (SMEs + institutions) with budgets for SaaS, but adoption hinges on trust in data privacy and accuracy. Option 1 (global/shared RAG + user-specific RAG) is superior for scalability and cost efficiency. Fine-tuning (Option 2) is overkill unless the domain is highly specialized (e.g., Sinhala legal jargon), which isn't specified. A curated global knowledge base ensures consistent, citable answers for common queries, while user-specific RAG handles private data. This hybrid approach leverages existing LLMs (reducing costs) and avoids the complexity of fine-tuning. Citations can be managed via retrieval-augmented pipelines (e.g., RAG with metadata linking to sources). Scalability is achievable with vector databases (e.g., Pinecone, Weaviate) and load-balanced LLM APIs. Key risks: Data privacy compliance (Sri Lankan laws like GDPR-equivalent clauses) and user trust in cloud storage. Mitigate with end-to-end encryption, zero-knowledge proofs, and SOC 2-like certifications. Competitors (e.g., DocuSign, local ERP tools) lack integrated RAG, creating a defensible niche. Pilot with 50-100 users in legal/finance sectors to validate demand before scaling.

Synthesized by meta/llama-4-maverick-17b-128e-instruct (fallback #1) · 2.8s