Hacker News .hnnew | past | comments | ask | show | jobs | submitlogin

Not really. The API these days tends to be JSON so you can just figure out how to it works and what represents what.

For example I've been able to reimplement xmltv scrapers for several sources in less than a 100 lines with Scrapy. It's not hard, just requires a little discretion.



The difficulty isn't in making a scraper for a single site, but rather in the general case.

That is, making a scraper that can be pointed at an arbitrary site not known at the time of development.


I assumed most people were talking about dealing with single sites. From your previous comment about API documentation and "hand evaluating" Javascript I gathered that you were too. How would those things help one solve the general case?




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: