TL

tigerless-labs/paper-radar

Security testing
95 stars 品質 70 トレンド 70

paper-radar What 28 AI labs published on arXiv, over any date range

概要

paper-radar What 28 AI labs published on arXiv, over any date range

README

paper-radar What 28 AI labs published on arXiv, over any date range

paper-radar finds the AI papers 28 labs put on arXiv, over any date range you pick. Every count splits two ways — papers the lab led, and papers it merely appears on.

Method: match against ROR and email domains, using the author block read off each paper’s HTML.

You need it because arXiv does not carry affiliation. The `` field has a 1% fill rate; Semantic Scholar reaches 7%; OpenAlex returns zero for preprints. So “what did Google publish in the last two weeks” has no API that answers it.

Install

Hand it to your agent

Paste this into Claude Code, Codex, or anything else with a shell:

Install the paper-radar skill from https://github.com/tigerless-labs/paper-radar

That is the whole install. The skill is one self-contained folder — skills/paper-radar/, pure stdlib, nothing to build — so the agent clones the repo and drops that folder wherever it keeps its skills. Say “for this project” if you want it local rather than global.

By hand

If you would rather do it yourself, or your agent has no shell:

git clone https://github.com/tigerless-labs/paper-radar
cp -r paper-radar/skills/paper-radar ~/.claude/skills/     # Claude Code
cp -r paper-radar/skills/paper-radar ~/.codex/skills/      # Codex
cp -r paper-radar/skills/paper-radar ~/.agents/skills/     # generic SKILL.md agents

Per-project instead of global: copy it into .claude/skills/ inside the repo.

Claude on the web, ChatGPT, and other upload-based agents

Download paper-radar-skill.zip from the latest release and upload it as a skill — Claude web takes it under Settings → Capabilities → Skills; ChatGPT and other hosted agents take the same folder as project files.

To build the zip yourself:

cd skills && zip -r ../paper-radar-skill.zip paper-radar

One caveat, stated plainly: the script fetches from arxiv.org over HTTPS. Hosted sandboxes that block outbound network will run the code and return nothing. This is a property of the sandbox, not of the skill — check whether yours allows network egress before relying on it there.

Use it

In Claude Code it is a slash command:

/paper-radar what did big tech publish on arXiv in the last two weeks?

In any agent, just ask — the skill’s description is written so it gets picked up on its own:

has Xiaomi published anything on self-evolving agents?
which companies led the most agent-memory papers this month?

Or run it yourself, no agent involved:

python3 skills/paper-radar/scripts/paper_radar.py --days 14

A two-week window (~3000 papers) takes about three minutes. Every flag is documented in the skill’s SKILL.md.

How it works

0 listing      arXiv API: category x submittedDate, deduped by ID
1 full text    arxiv.org/html/{id}, Range-request the first 90KB, cut ltx_authors to ltx_abstract
2 structure    map (author, superscripts) against (superscript, affiliation)
3 entity       email domain, then ROR name variants, then the lab alias table
4 grading      lead = the first author's institution; intern markers surfaced, not judged
5 supplements  Apple RSS, MSR embedded JSON, arXiv team-name queries
6 report       dedupe, write markdown, keep the evidence for every match

The report gives every lab two columns, and the gap between them is the whole point:

| Company         | Total hits | Of which lead |
|-----------------|-----------:|--------------:|
| Microsoft       |         31 |            18 |
| Meta            |          9 |             9 |
| Apple           |          8 |             0 |

The second column comes from the superscripts: whoever the first author is attached to led the paper. In one two-week window Google appeared on 10 papers and led 4; Adobe appeared on 5 and led 0.

Every match records what it matched on and the raw text it matched, so you can check any row by eye. Where it goes wrong, and what to do about it, is in the skill’s SKILL.md.

Coverage

Google · Google DeepMind · Microsoft · Meta · Apple · NVIDIA · Amazon · OpenAI · Anthropic · IBM · Adobe · Salesforce · Alibaba · Tencent · ByteDance · Baidu · Huawei · Xiaomi · Ant Group · Meituan · Kuaishou · DeepSeek · MiniMax · Moonshot AI · Zhipu AI · StepFun · Shanghai AI Lab · BAAI

26 companies plus two non-profit research institutes (Shanghai AI Lab, BAAI). Each is anchored to a real ROR ID, carrying 98 official name variants between them, Chinese included. Adding one is a single entry in scripts/data/ror_anchors.json — but anchor the ROR ID by hand. Fuzzy ROR search lands on the wrong entity often enough to matter: Adobe returns “Adobe Gastroenterology”, Meta returns “Meta House”, Google returns the Ireland subsidiary.

What it does not do

  • No fuzzy generalization. A company that isn’t in ror_anchors.json is invisible, full stop. This is the direct cost of using lookup instead of a model — in exchange, there are no confident wrong answers.
  • No heat signal. Measured: of 231 big tech papers in a two-week window, zero were discussed on Hacker News, while every HN-discussed paper in that window came from academia or a startup. Paper count is supply, not attention.
  • No PDF parsing. 12% of submissions ship as PDF only and are simply missing. Reading TeX source is the planned fix and is not built yet.
  • Per-company miss rates vary wildly. Apple’s authors usually omit their affiliation, so the author block alone misses ~78% of its papers; that one is patched from Apple’s official RSS feed. Do not apply a single confidence level to the whole report.

License

MIT

View this README on GitHub

推奨ツール

別のキーワードを試すか、フィルタを外してください。

インストール

npx skillfish add tigerless-labs/paper-radar