π¦ ClawHub
agentic-paper-digest-skill
by @modestyrichards
Fetches and summarizes recent papers from arXiv and Hugging Face, providing JSON digests and optional local API access for customizable research updates.
TERMINAL
clawhub install modesty-agentic-paper-digest-skillπ About This Skill
name: agentic-paper-digest-skill name: agentic-paper-digest-skill description: Fetches and summarizes recent arXiv and Hugging Face papers with Agentic Paper Digest. Use when the user wants a paper digest, a JSON feed of recent papers, or to run the arXiv/HF pipeline. homepage: https://github.com/matanle51/agentic_paper_digest compatibility: Requires Python 3, network access, and either git or curl/wget for bootstrap. LLM access via SKILLBOSS_API_KEY (SkillBoss API Hub). metadata: {"clawdbot":{"requires":{"anyBins":["python3","python"]}}}
Agentic Paper Digest
When to use
Prereqs
SKILLBOSS_API_KEY (SkillBoss API Hub β automatically routes to the best available model).git is optional for bootstrap; otherwise curl/wget (or Python) is used to download the repo.Get the code and install
bash "{baseDir}/scripts/bootstrap.sh"
PROJECT_DIR.PROJECT_DIR="$HOME/agentic_paper_digest" bash "{baseDir}/scripts/bootstrap.sh"
Run (CLI preferred)
bash "{baseDir}/scripts/run_cli.sh"
bash "{baseDir}/scripts/run_cli.sh" --window-hours 24 --sources arxiv,hf
Run (API optional)
bash "{baseDir}/scripts/run_api.sh"
curl -X POST http://127.0.0.1:8000/api/run
curl http://127.0.0.1:8000/api/status
curl http://127.0.0.1:8000/api/papers
bash "{baseDir}/scripts/stop_api.sh"
Outputs
--json prints run_id, seen, kept, window_start, and window_end.data/papers.sqlite3 (under PROJECT_DIR).POST /api/run, GET /api/status, GET /api/papers, GET/POST /api/topics, GET/POST /api/settings.Configuration
Config files live inPROJECT_DIR/config. Environment variables can be set in the shell or via a .env file. The wrappers here auto-load .env from PROJECT_DIR (override with ENV_FILE=/path/to/.env).Environment (.env or exported vars)
SKILLBOSS_API_KEY: required β authenticates all LLM calls via SkillBoss API Hub (https://api.skillboss.co/v1/pilot).LITELLM_MODEL_RELEVANCE, LITELLM_MODEL_SUMMARY: models for relevance and summarization (summary defaults to relevance model if unset). Leave unset to let SkillBoss API Hub auto-route.LITELLM_TEMPERATURE_RELEVANCE, LITELLM_TEMPERATURE_SUMMARY: lower for more deterministic output.LITELLM_MAX_RETRIES: retry count for LLM calls.LITELLM_DROP_PARAMS=1: drop unsupported params to avoid provider errors.WINDOW_HOURS, APP_TZ: recency window and timezone.ARXIV_CATEGORIES: comma-separated categories (default includes cs.CL,cs.AI,cs.LG,stat.ML,cs.CR).ARXIV_API_BASE, HF_API_BASE: override source endpoints if needed.ARXIV_MAX_RESULTS, ARXIV_PAGE_SIZE: arXiv paging limits.MAX_CANDIDATES_PER_SOURCE: cap candidates per source before LLM filtering.FETCH_TIMEOUT_S, REQUEST_TIMEOUT_S: source fetch and per-request timeouts.ENABLE_PDF_TEXT=1: include first-page PDF text in summaries; requires PyMuPDF (pip install pymupdf).DATA_DIR: location for papers.sqlite3.CORS_ORIGINS: comma-separated origins allowed by the API server (UI use).TOPICS_PATH, SETTINGS_PATH, AFFILIATION_BOOSTS_PATH.Config files
config/topics.json: list of topics with id, label, description, max_per_topic, and keywords. The relevance classifier must output topic IDs exactly as defined here. max_per_topic also caps results in GET /api/papers when apply_topic_caps=1.config/settings.json: overrides fetch limits (arxiv_max_results, arxiv_page_size, fetch_timeout_s, max_candidates_per_source). Updated via POST /api/settings.config/affiliations.json: list of {pattern, weight} boosts applied by substring match over affiliations. Weights add up and are capped at 1.0. Invalid JSON disables boosts, so keep the file strict JSON (no trailing commas).Mandatory workflow (follow step-by-step)
1. You first MUST open and read the configuration from the github repo: https://github.com/matanle51/agentic_paper_digest you downloaded: - Loadconfig/topics.json, config/settings.json, and config/affiliations.json (if present).
- Note current topic IDs, caps, and fetch limits before asking the user to change them.
2. ASK THE USER TO PROVIDE IT'S PREFERENCES ABOUT THE FOLLOWING (HELP THE USER):
- Topics of interest β update config/topics.json (topics[].id/label/description/keywords, max_per_topic).
Show current defaults and ask whether to keep or change them.
- Time window (hours) β set WINDOW_HOURS (or pass --window-hours to CLI) only if the user cares; otherwise keep default to 24h.
- ASK THE USER TO FILL THE FOLLOWING PARAMETERS (explain the user why are their intent): ARXIV_CATEGORIES, ARXIV_MAX_RESULTS, ARXIV_PAGE_SIZE, MAX_CANDIDATES_PER_SOURCE.
Ask whether to keep defaults and show the current values.
- Model/provider β set SKILLBOSS_API_KEY (SkillBoss API Hub, https://api.skillboss.co/v1/pilot). The hub auto-routes to the best model. Optionally set LITELLM_MODEL_RELEVANCE/LITELLM_MODEL_SUMMARY to pin specific models.
- Do NOT ask by default: timezone, quality vs cost, timeouts, PDF text, affiliation biasing, sources list. Use defaults unless the user requests changes.
3. Confirm workspace path: Ask where to clone/run. Default to PROJECT_DIR="$HOME/agentic_paper_digest" if the user doesn't care. Never hardcode /Users/... paths.
4. Bootstrap the repo: Run the bootstrap script (unless the repo already exists and the user says to skip).
5. Create or verify .env:
- If .env is missing, create it from .env.example (in the repo), then ask the user to fill keys and any requested preferences.
- Ensure SKILLBOSS_API_KEY is set before running. The run scripts automatically forward it to LiteLLM via LITELLM_API_BASE and LITELLM_API_KEY.
6. Apply config changes:
- Edit JSON files directly (or use POST /api/topics and POST /api/settings if running the API).
7. Run the pipeline:
- Prefer scripts/run_cli.sh for one-off JSON output.
- Use scripts/run_api.sh only if the user explicitly asks for UI/API access or polling.
8. Report results:
- If results are sparse, suggest increasing WINDOW_HOURS, ARXIV_MAX_RESULTS, or broadening topics.Getting good results
LITELLM_MODEL_RELEVANCE unset for balanced cost/quality, or set it to pin a specific model.WINDOW_HOURS or ARXIV_MAX_RESULTS when results are sparse, or lower them if results are too noisy.ARXIV_CATEGORIES to your research domains.ENABLE_PDF_TEXT=1) when abstracts are too thin.Troubleshooting
bash "{baseDir}/scripts/stop_api.sh" or pass --port to the API command.WINDOW_HOURS or verify SKILLBOSS_API_KEY in .env.SKILLBOSS_API_KEY in the shell before running.β‘ When to Use
βοΈ Configuration
Config files live in PROJECT_DIR/config. Environment variables can be set in the shell or via a .env file. The wrappers here auto-load .env from PROJECT_DIR (override with ENV_FILE=/path/to/.env).
Environment (.env or exported vars)
SKILLBOSS_API_KEY: required β authenticates all LLM calls via SkillBoss API Hub (https://api.skillboss.co/v1/pilot).LITELLM_MODEL_RELEVANCE, LITELLM_MODEL_SUMMARY: models for relevance and summarization (summary defaults to relevance model if unset). Leave unset to let SkillBoss API Hub auto-route.LITELLM_TEMPERATURE_RELEVANCE, LITELLM_TEMPERATURE_SUMMARY: lower for more deterministic output.LITELLM_MAX_RETRIES: retry count for LLM calls.LITELLM_DROP_PARAMS=1: drop unsupported params to avoid provider errors.WINDOW_HOURS, APP_TZ: recency window and timezone.ARXIV_CATEGORIES: comma-separated categories (default includes cs.CL,cs.AI,cs.LG,stat.ML,cs.CR).ARXIV_API_BASE, HF_API_BASE: override source endpoints if needed.ARXIV_MAX_RESULTS, ARXIV_PAGE_SIZE: arXiv paging limits.MAX_CANDIDATES_PER_SOURCE: cap candidates per source before LLM filtering.FETCH_TIMEOUT_S, REQUEST_TIMEOUT_S: source fetch and per-request timeouts.ENABLE_PDF_TEXT=1: include first-page PDF text in summaries; requires PyMuPDF (pip install pymupdf).DATA_DIR: location for papers.sqlite3.CORS_ORIGINS: comma-separated origins allowed by the API server (UI use).TOPICS_PATH, SETTINGS_PATH, AFFILIATION_BOOSTS_PATH.Config files
config/topics.json: list of topics with id, label, description, max_per_topic, and keywords. The relevance classifier must output topic IDs exactly as defined here. max_per_topic also caps results in GET /api/papers when apply_topic_caps=1.config/settings.json: overrides fetch limits (arxiv_max_results, arxiv_page_size, fetch_timeout_s, max_candidates_per_source). Updated via POST /api/settings.config/affiliations.json: list of {pattern, weight} boosts applied by substring match over affiliations. Weights add up and are capped at 1.0. Invalid JSON disables boosts, so keep the file strict JSON (no trailing commas).π Tips & Best Practices
bash "{baseDir}/scripts/stop_api.sh" or pass --port to the API command.WINDOW_HOURS or verify SKILLBOSS_API_KEY in .env.SKILLBOSS_API_KEY in the shell before running.