(Note: This is a fork of the original (now dead) warc repository. This is a rough python3 port / update that I used in my warc-extractor located here https://github.com/recrm/ArchiveTools. As the tool is complete, I have mostly stopped development of this library.)
WARC (Web ARChive) is a file format for storing web crawls.
This warc library makes it very easy to work with WARC files.:
import warc with warc.open("test.warc") as f: for record in f: print record['WARC-Target-URI'], record['Content-Length']
The documentation of the warc library is available at http://warc.readthedocs.org/.
This software is licensed under GPL v2. See LICENSE file for details.