@v-bible/bulac-scraper
v1.2.0
Published
Digital Bulac Library Scraper
Readme
:notebook_with_decorative_cover: Table of Contents
:star2: About the Project
:key: Environment Variables
To run this project, you will need to add the following environment variables to
your .env file:
- App configs:
LOG_LEVEL: Log level.LOG_FILE_PATH: (Optional) File path to save logs. Default toscraper.log.
E.g:
# .env
LOG_LEVEL=infoYou can also check out the file .env.example to see all required environment
variables.
:toolbox: Getting Started
:bangbang: Prerequisites
This project uses pnpm as package manager:
npm install --global pnpm
:running: Run Locally
Clone the project:
git clone https://github.com/v-bible/bulac-scraper.gitGo to the project directory:
cd bulac-scraperInstall dependencies:
pnpm installBuild the project:
pnpm build:eyes: Usage
[!NOTE] Support both "ark:" links (recommended) and manifest ("iiif") links. To get the manifest url of a document, you can go to the document page on Bulac, click on the "IIIF" button below the document viewer.
USAGE
bulac-scraper [--outDir value] [--height value] [--width value] [--toPdf] [--ignoreCompleted] [--overwrite] [--fromFile value] <args>...
bulac-scraper --help
bulac-scraper --version
Digital Bulac Library Scraper
FLAGS
[--outDir] Output directory. Default to "./output/<document-name>"
[--height] Image height. Default to 982 pixels
[--width] Image width
[--toPdf/--noToPdf] Convert downloaded images to a single PDF file
[--ignoreCompleted/--noIgnoreCompleted] Skip downloading if all images already exist in the output directory, or PDF already exists if --toPdf is set
[--overwrite/--noOverwrite] Overwrite existing files if they already exist in the output directory
[--fromFile] Path to a text file containing a list of document urls to scrape from Bulac (one url per line)
-h --help Print help information and exit
-v --version Print version information and exit
ARGUMENTS
args... List of document urls to scrape from Bulac (e.g., "https://bina.bulac.fr/s/bina/ark:/73193/bcrk5b", "https://bina.bulac.fr/iiif/2/579892/manifest")[!NOTE] For image size, you should only specify either height or width, but not both. This is because the images are served in a way that maintains their aspect ratio. If you specify both height and width, the images may be distorted. If you want to specify both, make sure you know the original aspect ratio of the images.
Example:
pnpm build && ./dist/cli.mjs --outDir ./my-output --toPdf https://bina.bulac.fr/s/bina/ark:/73193/bcrk5b https://bina.bulac.fr/iiif/2/572900/manifest
pmpn build && ./dist/cli.mjs --outDir ./my-output --toPdf --ignoreCompleted --overwrite --fromFile ./document-urls.txtCrawl URL from Bulac category
A small script is provided to crawl all document urls from a given Bulac
category page. You can run it as follows, requires
uv tool to be installed:
# scraper.py
# /// script
# dependencies = [
# "beautifulsoup4",
# "requests",
# ]
# ///
import time
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
def scrape_bina_ark_urls(start_url):
session = requests.Session()
session.headers.update(
{
"User-Agent": (
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
"AppleWebKit/537.36 (KHTML, like Gecko) "
"Chrome/120.0.0.0 Safari/537.36"
)
}
)
current_url = start_url
ark_urls = set()
page = 1
while current_url:
print(f"Scraping page {page}: {current_url}")
response = session.get(current_url)
if response.status_code != 200:
print(f"Failed to fetch page {page}. Status: {response.status_code}")
break
soup = BeautifulSoup(response.text, "html.parser")
# Extract links containing the ARK identifier pattern ('ark:/')
for a_tag in soup.find_all("a", href=True):
href = a_tag["href"]
if "ark:/" in href:
full_url = urljoin(current_url, href)
# Strip query strings and fragment identifiers
clean_url = full_url.split("?")[0].split("#")[0]
ark_urls.add(clean_url)
# Locate pagination "Next" link
next_link = soup.find("a", attrs={"rel": "next"}) or soup.find(
"a", class_=lambda c: c and "next" in c.split()
)
if next_link and next_link.get("href"):
next_href = next_link["href"]
current_url = urljoin(current_url, next_href)
page += 1
time.sleep(1)
else:
print("No next page link found. Pagination complete.")
current_url = None
return list(ark_urls)
if __name__ == "__main__":
target_url = (
"https://bina.bulac.fr/s/bina/item?Search=&property%5B0%5D%5Bproperty%5D=51&property%5B0%5D%5Btype%5D=eq&property%5B0%5D%5Btext%5D=https://www.idref.fr/029517486"
)
extracted_links = scrape_bina_ark_urls(target_url)
print(f"\nSuccessfully collected {len(extracted_links)} ARK URLs:")
for link in extracted_links:
print(link)uv run ./scraper.py:wave: Contributing
Contributions are always welcome!
Please read the contribution guidelines.
:scroll: Code of Conduct
Please read the Code of Conduct.
:warning: License
This project is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) License.
See the LICENSE.md file for full details.
:handshake: Contact
Duong Vinh - @duckymomo20012 - [email protected]
Project Link: https://github.com/v-bible/bulac-scraper.

