Tjatse/node-readability

Stars
341
Rank 123,998 (Top 3 %)
Language
JavaScript
Created over 10 years ago
Updated over 6 years ago

Tjatse/node-readability

Tjatse

There are no reviews yet. Be the first to send feedback to the community and the maintainers!

Scrape/Crawl article from any site automatically. Make any web page readable, no matter Chinese or English.

read-art

Readability reference to Arc90's.
Scrape article from any page (automatically).
Make any web page readable, no matter Chinese or English.

快速抓取网页文章标题和内容，适合node.js爬虫使用，服务于ElasticSearch。

Guide

How it works

In my case, the speed of spider is about 1500k documents per day, and the maximize crawling speed is 1.2k /minute, avg 1k /minute, the memory cost are about 200 MB on each spider kernel, and the accuracy is about 90%, the rest 10% can be fixed by customizing Score Rules or Selectors. it's better than any other readability modules.

(4) Server infos:

20M bandwidth of fibre-optical

8 Intel(R) Xeon(R) CPU E5-2650 v2 @ 2.60GHz cpus

32G memory

pm2-gui

An elegant web & terminal interface for Unitech/PM2.

ansi-html

An elegant lib that converts the chalked text to HTML.

pm2-ant

🐜 Unitech/PM2 performance monioring using Statsd and Graphite

spider2

A 2nd generation spider to crawl any article site, automatic read title and article.

req-fast

Fastest way to fetch the web content(HTML stream) from server, supports:redirects, auto decode(e.g.:Chinese), gzip, cookie, proxy...

TiToast

Android-like toast for titanium(Alloy)

TaurusAir

node-subs

Tiny elegant, chainable, fastest literal substitution.

dynamic-timer

Schedule execution of a one-time callback after delay milliseconds, automatic, intelligence and without bothering, the delay is calculated from different algorithms, e.g.: Lucas Sequence, Fibonacci Sequence, DaYan Series and Arithmetic Procession.

range

Range parser that parse range from string, e.g. "0, 1, 7~8, 9-10, 100~105" -> [[0, 1], [7, 10], [100, 105]]