01
Different data formats
RSS and GraphQL expose different field structures.
Automation & Digital Systems
Designed and implemented a scheduled research pipeline that collects AI developments from RSS and API sources, normalizes them into a shared schema, filters and deduplicates records, uses Gemini for structured relevance analysis, and stores useful findings in a searchable Notion research library.
Converted repetitive AI research into a scheduled, stateful pipeline that continuously discovers, evaluates, deduplicates, and organizes useful developments.
Keeping up with AI developments means monitoring several sources, separating useful updates from noise, avoiding repeated content, and organizing findings for later reference. Doing that manually quickly becomes repetitive.
I built an n8n pipeline that collects new items, converts different source formats into one structure, prioritizes what should be evaluated, uses Gemini for structured analysis, and stores relevant findings in Notion.
The goal was a small reliable information system: aware of completed work, conservative with API quota, resistant to duplicates, and capable of retrying after external-service failures.
Implementation scope
A typical research session required checking multiple sources, opening individual posts, judging AI relevance, writing summaries, and transferring notes into a database.
The automation therefore needed to solve more than aggregation. It had to decide what deserved AI processing, what had already completed, and what should happen after partial failure.
01
RSS and GraphQL expose different field structures.
02
Previously reviewed URLs can appear in later runs.
03
Not every collected item deserves an LLM request.
04
Repeated writes reduce trust in the research library.
05
Incomplete work must remain eligible for a later retry.
Design priorities
One downstream workflow for fundamentally different sources.
Filter, deduplicate, and prioritize before Gemini.
Remember only URLs that reached a completed outcome.
Protect both AI capacity and the final Notion database.
Leave incomplete work eligible for a future run.
Source-specific work happens at the edge: TechCrunch enters through RSS, while Product Hunt is retrieved through GraphQL and pre-filtered for AI signals.
After both sources are normalized and merged, every item follows the same state check, prioritization, AI analysis, duplicate protection, persistence, and terminal logging path.
Multi-source AI research pipeline architecture
Scheduled Trigger
Starts a controlled research run
Source-specific ingestion
TechCrunch AI
RSS feed
RSS Transformation
Maps feed fields to the shared schema
Product Hunt
Authenticated GraphQL API
AI Pre-Filter
Removes obviously unrelated products
API Transformation
Maps post fields to the shared schema
Data Normalization
One source-neutral research shape
Merge Research Sources
One downstream item stream
Persistent Processing Check
Removes URLs with completed outcomes
Sort by Publication Date
Newest eligible items first
Daily Processing Limit
Maximum five items per run
Gemini AI Analysis
Structured relevance and research context
Structured Research Record
Metadata and AI output combined
Relevant Research?
Routes the terminal outcome
Processing Log
Records the irrelevant terminal outcome
Notion Duplicate Check
Protects the destination independently
Processing Log
Records already_in_notion
Save to Notion
Creates the research page
Processing Log
Records saved_to_notion
Five stages turn mixed source data into consistent, searchable research.
A scheduled n8n run retrieves TechCrunch AI through RSS and Product Hunt posts through authenticated GraphQL requests.
A lightweight keyword check removes obviously unrelated Product Hunt posts before any limited AI quota is used.
Cheap deterministic filtering happens before expensive AI reasoning.
Both source formats are mapped to the same source-neutral fields, so downstream nodes do not need source-specific logic.
sourcesource_typetitleurlpublished_atauthorraw_excerptCompleted URLs are removed, eligible records are sorted newest first, and only the first five proceed to analysis.
Gemini returns structured analysis; relevant, genuinely new items reach Notion, and every completed path records a terminal outcome.
The native diagram explains the architecture; these screenshots show the published workflow and the data it produces.




Development used a relatively small Gemini free-tier request allowance, so quota became part of the processing design rather than an afterthought.
Completed records are removed first. The remaining items are sorted newest to oldest, capped at five, and processed individually with a delay and automatic retries for temporary failures.
Designed sequence
Wasteful sequence avoided
Gemini evaluates each eligible item and returns a constrained structure rather than unrestricted prose. Later nodes can route and store the response without interpreting free-form text.
{
"relevant": boolean,
"summary": string,
"why_it_matters": string,
"who_its_for": string,
"category": enum,
"importance": 1–10,
}
Allowed categories
A transformation node merges the source metadata with Gemini’s response. This source-neutral record is the contract used by both Notion and the processing log.
sourcesource_typetitleurlpublished_atauthorrelevantsummarywhy_it_matterswho_its_forcategoryimportanceThe workflow distinguishes between an item it has seen and one it has successfully finished processing. A URL enters the persistent processing log only after a terminal outcome.
Completed outcomes
Irrelevant items are still completed work. Relevant items receive a second URL check inside Notion, so the final library stays clean even if processing history and destination state ever diverge.
Before Gemini: “Has this URL already completed?”
Avoids repeated AI processing across scheduled runs.
Before storage: “Does this URL already exist?”
Protects destination integrity independently.
“Seen” and “successfully finished” are different workflow states. Preserving that distinction keeps failed work recoverable.
A scheduled workflow collects new research without repeated manual source checking.
RSS and authenticated API data enter the same normalized pipeline.
Deterministic filtering and freshness ranking reserve Gemini for eligible work.
Every accepted item reaches Notion with the same structured research fields.
Processing history and a destination lookup protect two different failure points.
Terminal logging supports repeated schedules and safe retries after incomplete runs.
Connecting feeds and APIs to an LLM was straightforward. The more valuable work was deciding when an item becomes complete, how different formats become interchangeable, where duplicate protection belongs, and what should happen when processing stops halfway through.
Those decisions transformed a chain of integrations into a small, resilient information pipeline.