Getting Started
Installation
You can install PageSieve from the Mozilla Add-on Store or by downloading the latest release from GitHub.
NoteBrowser Compatibility
The browser extension is currently built for Mozilla Firefox using WebExtensions Manifest v2. Support for Chrome and Chromium-based browsers is planned on the Roadmap.
Development Build
To build the extension from source:
- Clone the repository
- Install dependencies:
bun install - Build the extension:
just build-extension(orbun run build:extension) - In Firefox, go to
about:debugging-> This Firefox -> Load Temporary Add-on… and selectapps/extension/dist/manifest.json.
Extension Usage
- Open the PageSieve sidebar from your browser’s extension menu.
- Define Fields: Add field names for the data you want to extract.
- Select Elements: Use the point-and-click selector to identify the elements on the page.
- Pagination: Configure how the extension should navigate to the next page (optional).
- Extract: Click the “Scrape” button to start the process.
- Export: Download your results in JSON, CSV, or other supported formats.
Command Line Interface
Saved recipes can also be run headlessly using the CLI with either Cheerio (fast HTTP) or Playwright (full browser).
CLI Installation
Install globally via npm:
npm install -g pagesieve-cliOr build from the repository:
just install-cliIf using the Playwright engine, install Chromium:
npx playwright install chromiumCLI Basic Usage
Run a scrape recipe:
pagesieve run -c recipe.json -o results -f jsonVerify a recipe against the schema:
pagesieve verify -c recipe.json