Home · Data
Get to know the details.
Data, sources and our crawler
Every fact on this site carries its source, its license basis and the date it was last checked. This page lists where data comes from and what we never do.
Where the data comes from
| Tier | Source | License basis |
|---|---|---|
| 1 | Open data, public records and published award lists | cdla-permissive-2.0, apache-2.0, odbl-1.0, public-record, official-calendar, official-listing |
| 2 | The business's own website | robots-permitted-page |
| 3 | Licensed opinion source | licensed:<contract>, official-api:<name> |
| 4 | First-party data | first-party |
| 5 | Stated by the business | business-supplied |
What we never do
- We do not scrape Google, Yelp, Tripadvisor, Booking, Expedia or any site whose terms forbid it.
- We do not bypass blocks, rate limits, logins or CAPTCHAs, and we do not disguise our crawler.
- We do not store or republish review text. We store counted aspects, sentiment and dates.
- We do not use photos from brand sites, travel sites or Google. Photos come from the business with rights.
- We do not let Google review text or any copied star rating feed a score.
- We do not sell inclusion, rank or score. Sponsored placements are labelled and never change a score.
Download the data
Every hotel, sourced fact and ranking on this site, free to reuse under CC BY 4.0 with credit to Journeyed.
Our crawler
User agent: Mozilla/5.0 (compatible; JourneyedBot/1.0; +https://journeyed.org/data/crawler/)
- It identifies itself with the user agent above and obeys robots.txt, including Crawl-delay.
- It waits at least 10 seconds between requests to one site, fetches one page at a time, and reads at most 12 pages per site per refresh.
- It stops at a block, a login or a challenge page and does not retry with another agent, address or browser.
- To ask us to stop, block the crawler in robots.txt or write to hello@journeyed.org.
Sources
- Overture Maps Places · cdla-permissive-2.0 · checked
Last checked