SoPaper, So Easy
This is a project designed for researchers to conveniently access papers they need.
The command line tool sopaper
can automatically search and download paper
from Internet, given the title.
The downloaded paper will thus have a readable file name
(I wrote it at the beginning because I'm tired of seeing the file name being random strings).
It mainly supports searching papers in computer science.
How to Use
Install command line dependencies:
- pdftk command line executable.
- Using pdftk on OSX10.11 might lead to hangs. See here for more info.
- poppler-utils (optional)
Install python package:
pip install --user sopaper
Usage:
$ sopaper --help
$ sopaper "Distinctive image features from scale-invariant keypoints"
$ sopaper "https://arxiv.org/abs/1606.06160"
NOTE: If you are not in school, you may need proxy by environment variable http_proxy
and https_proxy
,
to be able to download from certain sites (such as 'dl.acm.org').
Features
The searcher
module will fuzzy search and analyse results in
- Google Scholar
and the fetcher
module will further analyse the results and download papers from the following possible sources:
- direct pdf link
- dl.acm.org
- ieeexplore.ieee.org
- arxiv.org
Searcher
and Fetcher
are extensible to support more websites.
The command line tool will directly download the paper with a clean filename.
All downloaded paper will be compressed using ps2pdf
from poppler-utils, if available.
TODO
- Fetcher dedup: when arxiv abs/pdf apperas both in search results, page would be downloaded twice (maybe add a cache for requests)
- Don't trust arxiv link from google scholar
- Is title correctly updated for dlacm?
- Extract title from bibtex -- more accurate?
- Fetcher for other sites