Metadata-Version: 2.4
Name: berrycrawl
Version: 0.2.0
Summary: Official Python SDK for the Berrycrawl API
Project-URL: Documentation, https://docs.berrycrawl.com
Project-URL: Repository, https://github.com/strawberry-labs/berrycrawl-python
Project-URL: Issues, https://github.com/strawberry-labs/berrycrawl-python/issues
Author: Strawberry Labs LLC
License-Expression: Apache-2.0
License-File: LICENSE
Requires-Python: >=3.9
Requires-Dist: anyio<5,>=4
Requires-Dist: httpx<1,>=0.27
Requires-Dist: pydantic<3,>=1.10
Requires-Dist: typing-extensions>=4.7
Description-Content-Type: text/markdown

# Berrycrawl Python SDK

[![pypi](https://img.shields.io/pypi/v/berrycrawl)](https://pypi.python.org/pypi/berrycrawl)

The official Python SDK for scraping, crawling, searching, mapping, structured extraction, screenshots, and brand profiles.

[Documentation](https://docs.berrycrawl.com) · [Dashboard](https://berrycrawl.com/app) · [GitHub](https://github.com/strawberry-labs/berrycrawl-python)

## Table of Contents

- [Installation](#installation)
- [Reference](#reference)
- [Usage](#usage)
- [Environments](#environments)
- [Async Client](#async-client)
- [Exception Handling](#exception-handling)
- [Advanced](#advanced)
  - [Access Raw Response Data](#access-raw-response-data)
  - [Retries](#retries)
  - [Timeouts](#timeouts)
  - [Custom Client](#custom-client)
- [Contributing](#contributing)

## Installation

```sh
pip install berrycrawl
```

## Reference

A full reference for this library is available [here](./reference.md).

## Usage

Set `BERRYCRAWL_API_KEY` to an API key from the [Berrycrawl dashboard](https://berrycrawl.com/app).

```python
import os
from berrycrawl import Berrycrawl

client = Berrycrawl(api_key=os.environ["BERRYCRAWL_API_KEY"])
page = client.scrape(url="https://example.com/pricing")
print(page.data["markdown"])
```

### Crawl a website

```python
job = client.crawl(url="https://example.com/docs", limit=50)
print(job.id)
```

### Search and map

```python
results = client.search(query="best headless browser libraries", limit=10)
site_map = client.map_(url="https://example.com", search="documentation")
```

### Retrieve a brand profile

```python
brand = client.brand.retrieve(url="https://stripe.com")
print(brand.data)
```

### Brand design system

Brand responses include an optional `branding` object for compatibility with older API deployments. When available, it contains the rendered light/dark scheme, semantic colors, typography, spacing, representative input and button styles, and semantic image roles. Use `branding.images.favicon` for the square icon and `branding.images.logo` for the wordmark.

## Environments

This SDK allows you to configure different environments for API requests.

```python
from berrycrawl import Berrycrawl
from berrycrawl.environment import BerrycrawlEnvironment

client = Berrycrawl(
    environment=BerrycrawlEnvironment.PRODUCTION,
)
```

## Async Client

The SDK also exports an `async` client so that you can make non-blocking calls to our API. Note that if you are constructing an Async httpx client class to pass into this client, use `httpx.AsyncClient()` instead of `httpx.Client()` (e.g. for the `httpx_client` parameter of this client).

```python
import asyncio

from berrycrawl import AsyncBerrycrawl

client = AsyncBerrycrawl(
    api_key="<token>",
)


async def main() -> None:
    await client.brand.retrieve(
        url="https://stripe.com",
    )


asyncio.run(main())
```

## Exception Handling

When the API returns a non-success status code (4xx or 5xx response), a subclass of the following error
will be thrown.

```python
from berrycrawl.core.api_error import ApiError

try:
    client.brand.retrieve(...)
except ApiError as e:
    print(e.status_code)
    print(e.body)
```

## Advanced

### Access Raw Response Data

The SDK provides access to raw response data, including headers, through the `.with_raw_response` property.
The `.with_raw_response` property returns a "raw" client that can be used to access the `.headers` and `.data` attributes.

```python
from berrycrawl import Berrycrawl

client = Berrycrawl(...)
response = client.brand.with_raw_response.retrieve(...)
print(response.headers)  # access the response headers
print(response.status_code)  # access the response status code
print(response.data)  # access the underlying object
```

### Retries

The SDK is instrumented with automatic retries with exponential backoff. A request will be retried as long
as the request is deemed retryable and the number of retry attempts has not grown larger than the configured
retry limit (default: 2).

Which status codes are retried depends on the `retryStatusCodes` generator configuration:

**`legacy`** (current default): retries on
- [408](https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/408) (Timeout)
- [409](https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/409) (Conflict)
- [429](https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/429) (Too Many Requests)
- [5XX](https://developer.mozilla.org/en-US/docs/Web/HTTP/Status#server_error_responses) (All server errors, including 500)

**`recommended`**: retries on
- [408](https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/408) (Timeout)
- [409](https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/409) (Conflict)
- [429](https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/429) (Too Many Requests)
- [502](https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/502) (Bad Gateway)
- [503](https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/503) (Service Unavailable)
- [504](https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/504) (Gateway Timeout)

Use the `max_retries` request option to configure this behavior.

```python
client.brand.retrieve(..., request_options={
    "max_retries": 1
})
```

### Timeouts

The SDK defaults to a 60 second timeout. You can configure this with a timeout option at the client or request level.

```python
from berrycrawl import Berrycrawl

client = Berrycrawl(..., timeout=20.0)

# Override timeout for a specific method
client.brand.retrieve(..., request_options={
    "timeout": 1
})
```

### Custom Client

You can override the `httpx` client to customize it for your use-case. Some common use-cases include support for proxies
and transports.

```python
import httpx
from berrycrawl import Berrycrawl

client = Berrycrawl(
    ...,
    httpx_client=httpx.Client(
        proxy="http://my.test.proxy.example.com",
        transport=httpx.HTTPTransport(local_address="0.0.0.0"),
    ),
)
```

## Contributing

While we value open-source contributions to this SDK, this library is generated programmatically.
Additions made directly to this library would have to be moved over to our generation code,
otherwise they would be overwritten upon the next generated release. Feel free to open a PR as
a proof of concept, but know that we will not be able to merge it as-is. We suggest opening
an issue first to discuss with us!

On the other hand, contributions to the README are always very welcome!
