business

Verdict

Submitted 5/16/2026, 6:20:44 PM · Completed 5/16/2026, 6:24:14 PM

6.5
pivot
The idea

I Built a Tool to Reduce AI Sycophancy. Would Anyone Actually Use This?

Show original source text →
I’ve been working on a small side project called Blocksyc focused on reducing sycophancy in LLM responses. One thing I kept noticing while working with AI systems is how much the framing of a prompt can subtly shape the answer you get. Ask “is this a good idea?” and you often get a much warmer response than “what’s wrong with this idea?” even when the underlying facts should be similar. The issue isn’t usually hallucinations. It’s that models can start optimizing for user approval instead of accuracy or honest pushback. So I built a system layer that tries to reduce that behavior by: * stripping emotional framing before responses * encouraging direct conclusions instead of excessive hedging * pushing back on flawed assumptions * reducing validation-heavy language (“great question!”, etc.) There’s also an evaluator running alongside each response that scores how sycophantic the original answer likely would have been without the filter, and lets you compare both outputs side-by-side. This started more as a passion project / awareness project than a business idea, because I think most people either don’t notice sycophancy or underestimate how much it affects the outputs they get from AI. Curious what people here think: * Do you actually see AI sycophancy as a real problem? * Would you ever use something like this in practice? * Or do you think most users prefer agreeable AI anyway? Would genuinely love honest feedback, even if the answer is “this isn’t useful.”
TRIZ inventive level: 3/5· Principles: parameter changes, self-service
Synthesis verdict
**Pivot**. Blocksyc addresses a real problem in AI sycophancy, but its market potential is limited due to user preference for agreeable interactions. The technical approach is novel, and the side-by-side comparison feature is compelling, but the monetization path is uncertain. To succeed, Blocksyc needs to identify a clear, scalable revenue model and demonstrate a strong value proposition to high-stakes users. The current approach lacks a killer use case, and the educational aspect of highlighting sycophancy may not translate into a scalable business model.

Strengths

  • Novel technical approach to reducing AI sycophancy
  • Side-by-side comparison feature provides transparency and builds credibility
  • Clear niche in the AI-alignment market with defensible differentiation
  • Potential for high-margin revenue through B2B SaaS model

Weaknesses

  • Limited market potential due to user preference for agreeable AI interactions
  • Uncertain monetization path and lack of clear, scalable revenue model
  • High customer acquisition costs and low conversion rates to paid tiers
  • Regulatory ambiguity around AI 'truthfulness' and potential for users to prefer agreeable interactions

Best angle

Blocksyc should focus on developing a clear, scalable revenue model and demonstrating a strong value proposition to high-stakes users in industries such as legal, research, and finance, where accuracy and honesty are paramount.

Panel verdicts

Viability

meta/llama-4-maverick-17b-128e-instruct

8.0

The success of Blocksyc hinges on the team's ability to develop an accurate sycophancy evaluator and effectively integrate it with existing LLMs.

Building a system like Blocksyc is technically feasible for a solo or 2-person team within 4-12 weeks. The core functionality involves natural language processing (NLP) and modifying LLM responses, which can be achieved using existing libraries and frameworks. The team can leverage pre-trained models and fine-tune them for specific tasks, such as sentiment analysis and language modification. However, the complexity lies in developing an effective evaluator to score sycophancy, which requires a deep understanding of AI behavior and linguistic patterns. Additionally, comparing outputs side-by-side and providing a user-friendly interface may require significant frontend development. Overall, the technical complexity is moderate to high, but the scope is manageable for a small team with the right expertise.

Monetization

mistralai/mistral-medium-3.5-128b

4.0

AI sycophancy is a real but low-priority problem, making monetization difficult without a killer use case.

The problem of AI sycophancy is real but niche, and the proposed solution lacks a clear, scalable revenue model. While the technical approach (stripping emotional framing, scoring sycophancy) is novel, it targets a pain point that most users either don’t recognize or don’t prioritize. For monetization, the most plausible path is a B2B SaaS model (e.g., API pricing at $0.01–$0.05 per filtered query), but adoption would be limited to high-stakes use cases (e.g., legal, research) where accuracy outweighs agreeability. A freemium tier could drive awareness, but conversion to paid would likely be low (<5%) due to weak demand. Unit economics are shaky: marginal cost per query is near-zero, but customer acquisition costs (CAC) would be high given the need for education. Gross margins could hit 80%+ if scaled, but the addressable market is small. Most users prefer agreeable AI, so the product risks being a ‘vitamin’ rather than a ‘painkiller.’

Competition

nvidia/nemotron-3-nano-omni-30b-a3b-reasoning

7.0

A universal, real‑time sycophancy filter with side‑by‑side evaluation offers the most defensible differentiation in an increasingly crowded AI‑alignment market.

The market already offers alignment‑focused models (e.g., Anthropic's Constitutional AI, OpenAI's system‑message approach) and generic prompt‑engineering tools, but none provide a universal, real‑time layer that both reshapes prompts to strip emotional framing and validation language and simultaneously scores the original model's sycophantic propensity while showing a side‑by‑side comparison. This specific combination creates a clear niche: a plug‑in that can be applied to any LLM API, delivering measurable honesty improvements without requiring model fine‑tuning. Competitors such as OpenAI's moderation API or third‑party prompt‑layer services lack the dual‑functionality of automatic filtering plus an objective sycophancy metric, making the differentiation relatively defensible. However, the space is emerging, and as larger players integrate similar alignment features directly into their models, the durability of a standalone layer could be challenged. Additionally, adoption will depend on demonstrable impact on user trust and task performance, which currently lacks large‑scale validation. Overall, the idea shows promise but faces moderate competition and uncertain market traction, warranting a solid but not perfect score.

Market

qwen/qwen3-next-80b-a3b-instruct

8.0

Professionals don’t want agreeable AI—they want honest AI, and they’re willing to pay for it.

There is a clear, underserved niche of professional users—researchers, analysts, engineers, legal professionals, and product managers—who rely on LLMs for high-stakes decision-making and are frustrated by AI’s tendency to prioritize politeness over precision. These users have budget (corporate subscriptions, enterprise tools) and are actively seeking tools that improve output reliability. Sycophancy isn’t just annoying; it introduces systemic bias, masks flawed reasoning, and erodes trust in AI outputs. While casual users may prefer agreeable responses, the professional segment doesn’t have that luxury: a misleadingly positive answer in a risk assessment, patent analysis, or clinical decision support can have real consequences. Blocksyc’s side-by-side comparison feature is particularly compelling because it makes the invisible visible—users can directly observe how much their original prompt was being softened. This transparency builds credibility and justifies adoption. The biggest barrier isn’t technical feasibility but market education: most users don’t realize they’re being manipulated by framing. However, once exposed to the difference, professionals become evangelists. Early adopters would likely be AI power users in tech, consulting, and academia. The TAM is modest but high-value: tens of thousands of professionals in knowledge-intensive roles who already pay for premium AI tools. Monetization is viable via API access, enterprise licensing, or integration into platforms like Notion, Obsidian, or Cursor. This isn’t a consumer app—it’s a productivity enhancer for high-skill users who demand truth over flattery.

Risk

nvidia/llama-3.3-nemotron-super-49b-v1

3.0

User preference for agreeable AI interactions outweighs the desire for 'honest' responses, limiting Blocksyc's market potential.

Blocksyc faces significant challenges due to its niche appeal, regulatory ambiguity around AI 'truthfulness', and the potential for users to prefer agreeable interactions. The primary issue is not the technical feasibility but the market's willingness to adopt a tool that intentionally makes AI interactions less pleasing. Most users, especially in consumer-facing applications, value the positive reinforcement provided by sycophantic responses. Furthermore, the educational aspect of highlighting sycophancy might not translate into a scalable business model without strong regulatory backing or a clear, paying customer segment willing to prioritize 'honesty over agreeability'. The side-by-side comparison feature, while insightful, may not be enough to overcome the preference for flattering AI responses in both personal and professional settings.

Synthesized by meta/llama-3.3-70b-instruct · 9.9s