business

Verdict

Submitted 5/25/2026, 11:28:02 AM · Completed 5/25/2026, 11:29:40 AM

5.5
pivot
The idea

How do I convert a Train Timetable into an Excel Document?

Pain point
Manual data entry of train timetables from PDFs is error-prone and time-consuming due to non-standard symbols and formatting.
Who has this problem
Transportation analysts and researchers needing to process multiple train timetables
Contradiction (TRIZ)
Wants accurate data extraction but cannot afford manual verification for large volumes
Ideal final result
Automated extraction of clean, structured timetable data with symbols properly categorized
Suggested solution
Use a combination of Power Query for initial data ingestion and an AI tool like Claude or Groq for symbol interpretation and data cleaning
Show original source text →
https://preview.redd.it/clc7x49sja3h1.png?width=1378&format=png&auto=webp&s=a458f48a67f509b345e3668e65612e917868afb2 https://preview.redd.it/zwaao59sja3h1.png?width=1359&format=png&auto=webp&s=c9156ad78c63e72b51ef55a667e2818d82f3fa1a Hello [r/excel](https://www.reddit.com/r/excel/) I'm a bit stuck with this one. I have a collection of timetable PDFs that I would like put onto a spreadsheet, so that I can do some research on the given data. The problem is: **Different Symbols on the timetable are making the problem more complex than it needs to be**. I've asked Gemini/Grok to develop programs which I can use python for (Novice), but no matter what, it cannot determine that the Blue Arrows are. The Blue Arrows usually just determine when a train is a slower service that gets bypassed by a faster service, but can also determine when a train splits/joins together. There's also another problem with those I have to solve but that's another thing. Then there's all the other symbols which create noise on my spreadsheet, such as the black boxed numbers next to the stations which denote interchange times. What I would like, is a way to extract these timetables into a clean, readable excel sheet that a program can parse through and extract data for me. **How do I do this?**
TRIZ inventive level: 3/5· Principles: parameter changes, mechanical interaction
Synthesis verdict
**Pivot**: The idea of building a tool to extract data from PDF timetables into a clean Excel sheet has a clear niche and potential for monetization. However, the technical challenges, particularly in accurately interpreting custom symbols, pose significant risks. The venture's success hinges on overcoming these challenges, which may require substantial development resources and expertise in PDF parsing and computer vision. While the market is willing to pay for such a tool, the competition from generic PDF-to-Excel tools and niche railway data services may reduce the venture's defensibility.

Strengths

  • Addresses a genuine pain point for a specific but valuable niche
  • Potential for high margins due to low incremental costs post-development
  • Willingness to pay exists among rail enthusiasts and transport professionals

Weaknesses

  • Technical challenges in accurately interpreting custom symbols
  • Competition from generic PDF-to-Excel tools and niche railway data services
  • Reliance on manual configuration for each timetable's unique layout and symbol set

Best angle

The venture should focus on developing a domain-specific parser that combines OCR with a symbol-recognition layer, automatically normalizing the timetable into a standardized schema, to differentiate itself from existing solutions.

Panel verdicts

Viability

meta/llama-4-maverick-17b-128e-instruct

6.0

The success of this project hinges on the ability to accurately identify and interpret the various symbols in the PDF timetables using a combination of PDF parsing and computer vision techniques.

The idea involves extracting data from PDF timetables into a clean Excel sheet. The task requires handling various symbols and complexities, such as blue arrows indicating slower services or train splits/joins, and black boxed numbers denoting interchange times. A solo or 2-person team can potentially build a v1 in 4-12 weeks using a combination of PDF parsing libraries (e.g., PyPDF2, pdfminer) and computer vision techniques (e.g., OpenCV) to identify and extract relevant data. However, the complexity lies in accurately identifying and interpreting the various symbols, which may require significant training data and fine-tuning of machine learning models. The team would need to have expertise in Python, PDF parsing, and computer vision. While it's feasible, the task is not trivial, and the team would need to overcome several technical challenges.

Risk

nvidia/llama-3.3-nemotron-super-49b-v1

2.0

Complex, non-standard PDF symbols and layouts make automated extraction impractical without significant, specialized development resources.

The proposed venture faces insurmountable challenges due to the complexity of PDF parsing with custom, non-standard symbols, lack of a scalable solution for symbol interpretation (e.g., Blue Arrows' dual meaning), and the reliance on manual configuration for each timetable's unique layout and symbol set. Automated tools (like Gemini/Grok) fail to interpret these nuances without extensive, bespoke programming, which a novice in Python cannot feasibly undertake. Moreover, the venture's success hinges on the accuracy of data extraction, which is compromised by the 'noise' from various symbols (e.g., black boxed numbers for interchange times), requiring labor-intensive cleaning. Given these technical hurdles and the niche problem it solves (tailored to a specific, potentially small user base of timetable analysts), the venture is unlikely to gain traction or become viable within a short timeframe.

Competition

nvidia/nemotron-3-nano-omni-30b-a3b-reasoning

6.0

A durable edge requires a symbol‑interpretation layer that generic PDF extractors lack, making the venture viable only if railway symbols stay stable and the parser can adapt to varied PDF layouts.

The core need—turning complex, symbol‑rich railway timetable PDFs into a clean, parseable Excel sheet—is clear, but the market already offers generic PDF‑to‑Excel tools (e.g., Tabula, Camelot) and a few niche railway data services that provide structured timetables via APIs. Those existing solutions extract raw text or tables but do not interpret railway‑specific symbols such as blue arrows that indicate slower services, bypasses, or split/join events. A new entrant could differentiate by building a domain‑specific parser that combines OCR with a symbol‑recognition layer, automatically normalizing the timetable into a standardized schema. This adds real value because users would no longer need to manually decode symbols or write custom scripts. However, durability is uncertain: the differentiation hinges on the stability of railway symbol conventions and the ability to handle diverse PDF layouts. If operators change symbols or introduce new visual encodings, the parser may require frequent updates, reducing long‑term defensibility. Additionally, the competitive landscape includes both generic OCR services (which can be extended) and specialized railway data providers that already supply clean data, meaning the entrant must either lock in a unique data source or continuously adapt to layout variations. Overall, the idea has a plausible niche and a clear differentiation point, but its durability depends on sustained symbol consistency and ongoing maintenance.

Monetization

mistralai/mistral-nemotron(fallback #1)

7.0

Accurate symbol interpretation is the critical factor that will determine the tool's market viability and pricing power.

The idea addresses a niche but valuable problem in data extraction from complex PDF timetables, which can be monetized through a SaaS model. The pricing could be tiered based on the volume of PDFs processed or the complexity of the symbols handled, ranging from $20/month for basic plans to $200/month for enterprise solutions. The conversion path involves offering a free trial to demonstrate the tool's effectiveness, followed by a subscription model. Unit economics would benefit from automation, reducing the cost-to-serve significantly. Margins could be high (70-80%) due to low incremental costs post-development. The key challenge is accurately interpreting the symbols, which, if solved, would make the tool indispensable for researchers and analysts in the transportation sector.

Market

moonshotai/kimi-k2.6(fallback #1)

7.0

A specialized PDF-to-structured-data tool for rail timetables captures a passionate, underserved niche where generic AI tools fail and users currently accept painful manual work.

This idea addresses a genuine pain point for a specific but valuable niche: rail enthusiasts, transport researchers, and professionals who need structured timetable data. The core audience includes: (1) hobbyist 'trainspotters' and railfans (estimated 100K+ globally with strong community engagement on forums like r/excel, r/trains), (2) transport planners and consultants who need historical timetable analysis, (3) academic researchers studying service patterns, and (4) open data advocates building public transport datasets. The unmet need is clear—PDF timetable extraction with semantic understanding of rail-specific symbols (arrows for splitting/joining, boxed numbers for interchange times) that generic OCR and LLM tools fail to parse. Current alternatives (manual entry, generic PDF parsers, AI tools like Gemini) all fail on the domain-specific symbols. Willingness to pay exists: rail enthusiasts spend significantly on niche tools (£50-200/year subscriptions common), and consultancy budgets for transport analysis run £500-2000/day. The market is constrained by niche size but amplified by data scarcity—structured historical timetable data is valuable and rarely available. A SaaS model targeting £20-50/month for enthusiasts and £200-500/month for professional tiers could work. The key risk is technical feasibility of symbol recognition accuracy; if solved, network effects from user-contributed symbol libraries could build defensibility. Comparable: RailMiles, RealTimeTrains, and OpenRailData show commercial viability in adjacent spaces.

Synthesized by meta/llama-3.3-70b-instruct · 27.2s