Metadata-Version: 2.4
Name: instagram-posts-scraper
Version: 0.2.0
Summary: Implement Instagram Posts Scraper for post data retrieval
Home-page: https://github.com/FaustRen/instagram-posts-scraper
Author: FaustRen
Author-email: faustren1z@gmail.com
License: MIT
Classifier: Programming Language :: Python :: 3.11
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: beautifulsoup4>=4.13.4
Requires-Dist: cloudscraper>=1.2.71
Requires-Dist: lxml>=6.1.1
Requires-Dist: pandas>=2.2.3
Requires-Dist: pytz>=2024.2
Requires-Dist: requests>=2.32.3
Requires-Dist: selenium>=4.33.0
Requires-Dist: seleniumbase>=4.39.2
Dynamic: author
Dynamic: author-email
Dynamic: classifier
Dynamic: description
Dynamic: description-content-type
Dynamic: home-page
Dynamic: license
Dynamic: license-file
Dynamic: requires-dist
Dynamic: requires-python
Dynamic: summary

## 🚀 Sponsored by CoreClaw

Looking for a production-ready Instagram scraping API?

CoreClaw provides APIs and open-source Workers for **Instagram posts, profiles, comments, and more**, helping developers collect structured data at scale.

🎁 **Start free** → [https://coreclaw.com](https://www.coreclaw.com/?utm_source=github&utm_medium=cpc&utm_campaign=ren&utm_term=&utm_id=ren)

# Instagram Posts Scraper

InstagramPostsScraper is a Python library for collect instagram users' data.

The data obtained by web crawlers is not real-time data, but rather data from a specific point in time on the same day.

I’d really appreciate your support! You can star ⭐ or fork this repository to help me keep sharing more interesting web scrapers.

# Support Me

If you enjoy this project and would like to support me, please consider donating 🙌  
Your support will help me continue developing this project and working on other exciting ideas!

## 💖 Ways to Support:

- **PayPal**: [https://www.paypal.me/faustren1z](https://www.paypal.me/faustren1z)
- **Buy Me a Coffee**: [https://buymeacoffee.com/faustren1z](https://buymeacoffee.com/faustren1z)

Thank you for your support!! 🎉


## Requirements
```bash
beautifulsoup4==4.13.4
cloudscraper==1.2.71
lxml==6.1.1
pandas==2.2.3
pytz==2024.2
requests==2.32.3
selenium==4.33.0
seleniumbase==4.39.2
```

## Installation

To install the latest release from PyPI:

```sh
pip install instagram-posts-scraper
```

## Usage - Sample

```python
from instagram_posts_scraper.instagram_posts_scraper import InstaPeriodScraper
from IPython.display import display

ig_posts_scraper = InstaPeriodScraper()
target_info = {"username": "stephencurry30", "days_limit": 30}
res = ig_posts_scraper.get_posts(target_info=target_info)
display(res)
```

### Optional parameters

- **username**: target instagram user 
- **days_limit**: Number of days within which to scrape posts..

## Version

You can check the installed version and module documentation:

```python
import instagram_posts_scraper

print(instagram_posts_scraper.__version__)  # e.g. 0.2.0
print(instagram_posts_scraper.__doc__)      # module documentation
```

## Sample Output

The scraper returns a single consolidated dictionary containing the target's
normalized `profile`, the `account_status`, the scraping timestamp
(`updated_at`), a `posts` list of normalized posts, plus the raw `init_posts`
(picnob first-page HTML posts) and `top_posts` (profile-scraper highlights)
collections, which are preserved verbatim so no source data is lost.

Profile metadata is normalized: `followers` comes from the profile scraper's
precise count, while `following` and the biography fallback come from picnob.
Each entry in `posts` is normalized to a single, consistent engagement shape
(`like_count` / `comment_count` as integers). `init_posts` and `top_posts` keep
their original shapes untouched.

Below is an **abbreviated** example (long media URLs are truncated with `...` for readability).
For the complete, real output see
[`examples/example_output.json`](examples/example_output.json).

```jsonc
{
  "profile": {
    "username": "stephencurry30",
    "userid": "324599988",
    "full_name": "Wardell Curry",
    "biography": "Believer. Husband. Father. Founder. Philanthropist. Olympic Gold Medalist. NYT Best Selling Author. Philippians 4:13.",
    "followers": 57049215,
    "following": 1296,
    "posts_count": 1556,
    "profile_picture": "https://cdn.iqsaved.com/..."
  },
  "account_status": "public",
  "updated_at": "2026-08-21 15:28:13.009746+08:00",
  "posts": [
    {
      "shortcode": "6772442523573164715722",
      "caption": "Played a lil G with my boy @stephencurry30 this week to kick off Father’s Day weekend! ...",
      "media_type": "igtv",
      "is_video": true,
      "timestamp": 1781966678,
      "like_count": 98397,
      "comment_count": 540,
      "thumbnail": "https://scontent-ord5-1.cdninstagram.com/...",
      "image_url": "https://scontent.cdninstagram.com/..."
    }
    // ... more posts
  ],
  "init_posts": [
    {
      "text": "Quality time looks a little different in our family ...",
      "likes": "101k",
      "comments": "425",
      "time": "7 days ago",
      "thumbnail": "https://sp1.pixnoy.com/..."
    }
    // ... picnob first-page posts, preserved verbatim (now includes the cover `thumbnail`)
  ],
  "top_posts": [
    {
      "timestamp": 1786636718,
      "caption": "Quality time looks a little different in our family 😂 ...",
      "comment_count": 425,
      "like_count": 100894,
      "shortcode": "Db_GagbB20Q"
    }
    // ... profile-scraper highlights, preserved verbatim
  ]
}
```
