On this page
Webz.io vs. Perigon: News API Comparison for Developers

Webz.io vs. Perigon: News API Comparison for Developers

Webz.io vs. Perigon: News API Comparison for Developers

Perigon and Webz.io both give developers structured access to global news. Both support Boolean search, semantic search, sentiment analysis, summaries, source filters, and MCP integrations.

The difference becomes clearer when you look beyond raw result counts.

In a live benchmark, Perigon returned more raw results for several broad English queries. But Perigon includes reprints by default. Across ten English searches, 66.4% of Perigon’s raw results disappeared when its own reprint filter was applied.

When the same independent duplicate-detection method was applied to 500 results from each API, Webz.io returned:

  • 435 distinct results, compared with 389 from Perigon
  • 37% more distinct English results
  • A 13.0% duplicate rate, compared with 22.2% for Perigon
  • More than four times the estimated distinct coverage for the query “data breach”

Webz.io also returned substantially more Spanish, Hebrew, and French content, while providing a dedicated News API that is separate from its Blogs API. Perigon’s documented source universe includes company blogs and press releases alongside news publishers.

Benchmark results at a glance

Metric Webz.io Perigon
Distinct results in 500-record sample 435 389
Distinct English results across three sampled queries 272 199
Duplicate rate in matched sample 13.0% 22.2%
Clearly flagged aggregators, mirrors, community or non-news pages 9.4% 17.6%
Estimated distinct “data breach” results 936 234
Spanish inteligencia artificial raw results 11,850 3,682
Hebrew ישראל raw results 943 343
Documented language filters 76 28
Fixed documented topics 629 IPTC topics Evolving topic list

The benchmark used a fixed 24-hour publication window covering September 3, 2026, from 00:00:00 through 23:59:59 UTC.

Thirteen matched queries were used for full-day coverage counts. Duplicate analysis used the newest 100 results from each provider for five representative searches, producing a 1,000-record sample.

Raw result counts do not equal unique coverage

Perigon provides useful reprint detection. Every article can include a reprint flag and reprintGroupId, and developers can remove recognized copies by setting:

showReprints=false

The issue is that Perigon includes those copies by default. Its documentation also says that reprint detection looks back three days, so copies published outside that window may not be identified.

This had a large effect on the benchmark.

Query Perigon raw results Perigon without recognized reprints Removed
OpenAI 4,794 1,873 60.9%
“artificial intelligence” 18,745 6,954 62.9%
Tesla 2,977 991 66.7%
Ukraine 8,528 3,099 63.7%
“climate change” 4,718 1,803 61.8%
inflation 14,204 4,819 66.1%
“data breach” 425 216 49.2%
football 32,934 11,625 64.7%
“Donald Trump” 19,470 4,437 77.2%
“quantum computing” 1,089 415 61.9%

Across these ten searches, Perigon reported 107,884 raw matches. Its own reprint filter reduced that figure to 36,232.

That does not mean the remaining 36,232 results were necessarily all unique. It means they were not classified as reprints by Perigon’s three-day detection process.

Webz.io also exposes syndication metadata, including whether a result is syndicated, a shared syndication identifier, and the first observed copy.

To avoid relying solely on either provider’s internal classifier, the benchmark applied the same URL normalization and near-identical-title grouping to both 500-record samples.

Webz.io returned more distinct English results

The clearest difference appeared in the three English queries included in the independent deduplication sample.

Query Webz.io distinct results Perigon distinct results
OpenAI 95 77
Ukraine 84 67
“data breach” 93 55
Total 272 199

Webz.io returned 36.7% more distinct English results across these searches.

The “data breach” query produced the largest gap. Webz.io returned 1,006 raw full-day results, compared with 425 from Perigon. In the sampled top 100:

  • 93 Webz.io results were distinct
  • 55 Perigon results were distinct
  • One Perigon cluster contained 28 versions of essentially the same Reuters report

Applying the observed distinct share to the full-day result count produces a directional estimate of:

  • 936 distinct Webz.io results
  • 234 distinct Perigon results

That is approximately a 4× difference for a cybersecurity-focused query.

Stronger multilingual coverage

Webz.io led the French, Spanish, and Hebrew coverage tests.

Query Language Webz.io Perigon Webz.io lead
Macron French 574 332 1.7×
inteligencia artificial Spanish 11,850 3,682 3.2×
ישראל Hebrew 943 343 2.7×

Perigon’s Spanish and Hebrew top-100 samples contained fewer duplicates than Webz.io’s corresponding samples. But Perigon’s underlying result pools were much smaller.

After applying the observed sample distinct rates, Webz.io still produced approximately:

  • 2.8× more distinct Spanish results
  • 2.3× more distinct Hebrew results

Webz.io documents 76 supported language filters. Perigon currently lists 28 supported languages and provides English translations for non-English articles.

For global media monitoring, regional risk analysis, multilingual adverse-media screening, and international market intelligence, this difference matters more than a higher raw count on a few broad English queries.

A cleaner boundary between news and blogs

Perigon describes its source coverage as including:

  • Major media publishers
  • Local news outlets
  • Trade publications
  • Company blogs
  • Press releases
  • Other publicly available sources

These sources are part of the same broader content universe. Perigon offers labels such as Non-news, Opinion, Paid News, Press Release, Low Content, and Synthetic, so developers can attempt to exclude unwanted content.

Webz.io takes a different approach. It provides separate endpoints for separate content types:

News API      → articles from news sources

Blogs API     → blog and self-published content

Forums API    → discussions, posts and replies

Reviews API   → reviews and ratings

The APIs use a similar query language and normalized response structure, but developers choose the content type before running the query.

This separation gives applications a cleaner starting point.

A media-monitoring platform searching the News API does not need to begin every request by excluding company blogs and press releases. A market-research application that wants those sources can query the Blogs API separately and decide how to combine the datasets.

The distinction is especially useful when different content types need different:

  • Ranking logic
  • Retention policies
  • Trust rules
  • User interfaces
  • Alerts
  • Pricing or usage limits

No open-web news index contains only perfect publishers. But separating news, blogs, forums, and reviews at the product level makes the source mix more predictable.

Source quality is not the same as source volume

The benchmark also reviewed the source domains behind 500 results from each provider.

The audit conservatively flagged only clearly identifiable aggregators, content mirrors, community sites, or obvious non-news pages. It was not intended to assign a full credibility score to every publisher.

Source-audit result Webz.io Perigon
Records reviewed 500 500
Clearly flagged records 47 88
Flagged share 9.4% 17.6%

Neither API returned only traditional newsrooms.

Webz.io’s sample included some community pages, content mirrors, automated sites, and low-quality domains. Perigon’s sample included substantial volumes from aggregators, regional Yahoo copies, portals, and local sites carrying syndicated wire stories.

The difference is that Webz.io provides several direct controls for narrowing news results by source quality and type.

Webz.io provides a deeper trust layer

Webz.io’s News API returns and filters on source-level trust metadata, including:

  • Trusted-news classification
  • Fake-news classification
  • Satirical-news classification
  • Top-news source lists
  • Political bias
  • Newsroom, local-news, and government-news source types
  • Domain rank
  • Source country
  • Licensing and AI-use metadata

It also supports filters such as:

trust.category:trusted_news

trust.top_news:top_news

trust.source.type:newsroom

domain_rank:<100000

country:US

The News API documentation lists source, trust, domain-rank, language, category, topic, entity, ticker, sentiment, breaking-news, and syndication filters.

Perigon has useful source controls of its own, including source groups, audience estimates, global source rank, political leaning, paywall metadata, and editorial labels. Its Non-news, Paid News, Press Release, and Synthetic labels are helpful when cleaning a broader source universe.

However, the reviewed Perigon documentation does not describe a direct equivalent to Webz.io’s trusted, fake, and satirical source classification.

For applications where source credibility is part of the query—not merely something evaluated after retrieval—Webz.io provides the more explicit model.

More predictable categories and topics

Both APIs categorize and tag content, but their taxonomies work differently.

Webz.io

Webz.io documents:

  • 17 top-level news categories
  • 629 IPTC-derived topics
  • IPTC level-two and level-three hierarchy
  • Categories and topics returned with each matching article
  • Direct category: and topic: filtering

Examples include:

category:”Science and Technology”

topic:”cyber crime”

topic:”climate change”

topic:terrorism

The fixed topic list makes it easier to build repeatable rules, saved searches, dashboards, and customer-facing filters.

Perigon

Perigon documents:

  • 15 broad categories
  • An evolving topic list
  • Google Content Categories
  • Editorial labels
  • Article and video media types

Its labels provide useful distinctions such as opinion, press release, paid news, fact check, roundup, low content, and synthetic content.

Perigon’s model is flexible. Webz.io’s model is easier to use when an application needs a known, stable taxonomy that can be mapped into production rules.

Sentiment beyond the article level

Both APIs provide article-level sentiment.

Perigon returns numeric positive, negative, and neutral scores that add up to one. This is useful when developers need adjustable thresholds rather than a single label.

Webz.io returns an overall positive, negative, or neutral classification. It also goes further by attaching sentiment to named entities:

{

  “organizations”: [

    {

      “name”: “Example Company”,

      “sentiment”: “negative”

    }

  ],

  “persons”: [

    {

      “name”: “Example Executive”,

      “sentiment”: “neutral”

    }

  ]

}

Entity-level sentiment is available for people, organizations, and locations. Organization entities can also include stock tickers and exchanges.

This matters because the overall tone of an article is not always the tone directed at the company or person being monitored.

A positive article about a growing market may contain negative news about one supplier. A negative article about an industry downturn may describe one company as outperforming its peers. Document-level sentiment alone cannot reliably represent those distinctions.

Summaries are included by default

Both providers return summaries with article results.

Webz.io’s standard News API schema includes:

  • Full article text
  • Summary
  • Author
  • Publication and crawl time
  • Language
  • Sentiment
  • Categories
  • Topics
  • Entities
  • Images
  • Source metadata
  • Trust metadata
  • Syndication data

In the 500-result samples:

Summary completeness Webz.io Perigon
Dedicated summary populated 73.8% 57.2%

Perigon supplied a description for almost every sampled article, even when its dedicated summary field was empty.

Perigon also offers a separate multi-article summarizer that can generate a custom overview from up to 100 articles or one article per cluster. That is a useful capability for digests and briefings.

For normal article retrieval, however, Webz.io already includes its summary field in the standard news response.

Boolean search, semantic search, and MCP

Webz.io is not limited to keyword search.

It provides two complementary news-retrieval models.

News API

The News API is designed for precise, repeatable monitoring queries using:

  • Keywords and phrases
  • AND, OR, and NOT
  • Field-level filters
  • Proximity operators
  • Source and domain constraints
  • Entities and tickers
  • Categories and topics
  • Sentiment
  • Trust
  • Time ranges

For example:

curl -G “https://api.webz.io/api/news” \

  –data-urlencode ‘token=YOUR_TOKEN’ \

  –data-urlencode ‘q=(“data breach” OR ransomware) language:english trust.category:trusted_news’

News Search API

The News Search API accepts natural-language queries and retrieves articles by meaning rather than only exact keyword matches. It returns matching articles and relevant text excerpts for AI grounding and research workflows.

Example queries include:

Find reports of European manufacturers cutting production because of weak demand.

Which software companies disclosed security incidents affecting customer data?

What recent developments could disrupt semiconductor supply chains?

MCP server

Webz.io also provides a hosted MCP server. Compatible clients receive a news_search_by_webz tool that can call News Search directly from Cursor, Claude, ChatGPT, or another MCP-compatible client.

Perigon also provides vector search and an MCP server, so semantic retrieval and agent connectivity are available from both providers.

The choice is not “semantic search versus no semantic search.” The difference is what sits underneath that search: coverage, source separation, trust controls, taxonomy, and the uniqueness of the returned articles.

Feature comparison

Search and coverage

Feature Webz.io Perigon
Keyword search Yes Yes
Boolean search Yes Yes
Field-level query syntax Extensive Yes
Natural-language search News Search API Vector Search
Hosted MCP server Yes Yes
News and blogs separated Yes, separate APIs Mixed source universe, filter with labels
Documented language filters 76 28
Categories 17 15
Topics 629 fixed IPTC topics Evolving topic list
Multilingual benchmark result Led French, Spanish and Hebrew Lower counts in tested languages
Cybersecurity benchmark 4× estimated distinct “data breach” coverage Lower coverage in test
Source requests UI and API on eligible plans Paid custom-source process
Large-result pagination Cursor-based Page-based, maximum 10,000 per query

Perigon documents a maximum paginated result set of 10,000 records per query and recommends splitting larger searches by date. Webz.io returns an opaque continuation cursor for walking through results over time.

Enrichment and quality controls

Feature Webz.io Perigon
Full article text Yes Yes
Summary in article response Yes Yes
Article sentiment Positive/negative/neutral Numeric three-dimensional score
Entity extraction Yes Yes
Entity-level sentiment Yes No direct equivalent found in reviewed docs
Company tickers Yes Yes
Source political bias Yes Yes
Domain/source rank Yes Yes
Trusted-news classification Yes No direct equivalent found
Fake-news classification Yes No direct equivalent found
Satirical-news classification Yes No direct equivalent found
Newsroom/local/government source types Yes Source metadata and groups
Editorial labels Trust and source classifications Non-news, opinion, paid, press release, synthetic and more
Syndication metadata Yes Yes
One-switch reprint suppression Limited by available syndication classification Yes
Story clustering No dedicated story endpoint in the News API Yes
Custom multi-article summarizer Not a separate News API endpoint Yes

Developer access and pricing

Feature Webz.io Perigon
Free access $5 in API credit every month 150 requests per month
Credit card required to start No No
Free-plan commercial use Usage-based Webz.io access Personal/non-commercial license
Pay-as-you-go option Yes Paid monthly plans
Listed entry paid plan Usage-based $250/month for eligible startups
MCP Yes Yes
Historical datasets Live history plus Archive and data dumps since 2008 Plan-dependent, up to 10+ years
Separate blogs/forums/reviews products Yes No equivalent content separation

Webz.io’s current Free plan includes $5 in monthly credits with no credit card required, followed by pay-as-you-go access. Perigon’s current Free plan includes 150 monthly requests under a personal-use license; its listed Basic plan starts at $250 per month for eligible startups.

Where Perigon is strong

Perigon has several good features.

Its showReprints=false option and reprint-group identifiers are convenient. Its Stories API clusters related reporting around an event. Its separate summarizer can create a briefing across multiple articles. It also provides journalist, company, people, source, and topic endpoints through its API and MCP server.

Those capabilities may suit teams that want an application-ready event and media database with clustering built in.

But those features do not replace source breadth, source separation, and distinct article coverage.

Why developers choose Webz.io

Webz.io is the stronger option when the application needs:

  • More distinct results rather than more syndicated copies
  • Broad multilingual and regional news coverage
  • Cybersecurity and adverse-media coverage
  • A news endpoint separated from blogs and other open-web content
  • Trusted, fake, and satirical source classifications
  • Newsroom, local-news, and government-source filters
  • Entity-level sentiment
  • A fixed IPTC-based topic taxonomy
  • Large-scale cursor pagination
  • Boolean monitoring and natural-language semantic search
  • MCP access for AI agents
  • Pay-as-you-go access without a large monthly commitment

Perigon’s raw volume can look larger on broad English queries. Once reprints and source mix are examined, the difference becomes much smaller—and in the matched duplicate analysis, Webz.io returned more distinct content.

Frequently asked questions

Does Perigon return duplicate articles?

Perigon includes reprints by default. Developers must add showReprints=false to remove articles recognized by its reprint system. Perigon’s documentation says that recognition uses a three-day comparison window, so later copies may not be marked.

In the benchmark, Perigon’s native reprint filter removed 66.4% of its raw English results.

Does Perigon include blogs in its news results?

Perigon’s source documentation says its monitored sources include company blogs, press releases, trade publications, local outlets, major media publishers, and other publicly available sources.

Perigon provides labels and source groups to help filter that content.

Webz.io instead offers dedicated News and Blogs APIs. This makes it easier to keep editorial news separate from company commentary, independent publishing, and other blog content.

Are all Webz.io results from major news organizations?

No open-web API should be treated as a whitelist of major publishers.

Webz.io covers large, local, regional, specialist, and niche sources. Its advantage is that developers can use domain rank, top-news lists, trusted-news classifications, source types, political bias, and explicit domain filters to define the source set they need.

Does Webz.io have semantic or vector-based news search?

Yes. Webz.io’s News Search API accepts natural-language queries and finds articles based on meaning rather than only exact keyword matches.

The standard News API remains available for precise Boolean and field-based monitoring.

Does Webz.io have an MCP integration?

Yes. Webz.io provides a hosted MCP server for its News Search API. It gives compatible AI clients a news_search_by_webz tool.

Does Webz.io return summaries?

Yes. The standard News API response includes a summary field alongside full text, sentiment, categories, topics, entities, source data, and trust metadata.

In the benchmark sample, Webz.io’s dedicated summary field was populated more often than Perigon’s.

Which API is better for media monitoring?

Webz.io is the better fit when the monitoring system needs broad coverage, multilingual sources, Boolean rules, source-quality filters, trust signals, entity sentiment, and deep pagination.

Perigon may be attractive when built-in story clustering and multi-article summarization are the main requirements.

Which API is better for cybersecurity and adverse media?

The benchmark strongly favored Webz.io. For “data breach”, Webz.io returned more than twice the raw coverage and approximately four times the estimated distinct coverage.

Webz.io’s entity, sentiment, source-country, trust, domain-rank, category, topic, and source-type fields also make it easier to build repeatable adverse-media and cyber-risk queries.

Which API is easier to test?

Both offer free access without requiring a credit card.

Webz.io provides $5 in free API credit every month and supports usage-based payment after the free allowance. Perigon’s free plan currently includes 150 requests per month under a personal-use license.

The bottom line

Perigon is a capable news intelligence API, particularly for story clustering, reprint grouping, and generated multi-article summaries.

But raw result volume does not tell the full story.

The duplicate-aware benchmark found that Webz.io returned:

  • 12% more distinct results overall
  • 37% more distinct English results
  • A 41% lower duplicate rate
  • Around 4× more distinct data-breach coverage
  • Far broader Spanish, Hebrew, and French coverage
  • Fewer clearly identifiable aggregators and non-news pages in the sampled results

Webz.io also provides something Perigon’s combined source universe does not: a clear separation between news, blogs, forums, and reviews.

For developers building media-monitoring systems, risk platforms, cybersecurity products, market-intelligence tools, or news-grounded AI applications, Webz.io provides the stronger combination of coverage, control, enrichment, and source transparency.

Start with $5 in free Webz.io API credit every month. No credit card required.

 

Subscribe to our blog for more news and updates!

By submitting you agree to Webz.io's Privacy Policy and further marketing communications.

Footer Background Large
Footer Background Small

Power AI with Web Data

icon

Ready to Explore Web Data at Scale?

Speak with a data expert to learn more about Webz.io’s solutions
Speak with a data expert to learn more about Webz.io’s solutions
Create your API account and get instant access to millions of web sources
Create your API account and get instant access to millions of web sources