Identify names in an article

Beta
Extract every company and person an article names and identify each one
View as Markdown

An article lookup reads one article, extracts every company and person it names, and identifies each name as POST /v1/lookup would, with the article as the source. It shares the lookup’s plan requirement.

  1. Start the job with POST /v1/lookup-jobs and a body holding sourceUrl, sourceNewsId, or both. The call returns 202 with a jobId. Sending the same article again soon returns the earlier job.
  2. Read the job with GET /v1/lookup-jobs/{jobId} until state is COMPLETED or FAILED. Only the caller that started the job can read it.
  3. Act on each entry in mention. Every entry has the name as the article writes it and a mentionType of COMPANY or PERSON:
    • identification is the lookup answer, read exactly as in Act on the answer.
    • shell is set when the answer is NO_MATCH and the job created a hidden placeholder record for the name. The job creates one only when the article and a live page on another site both name the subject. shell holds its entityId or personId; creating it costs nothing, and you can send it straight to a research update.
    • shellDetail says, in one sentence, why no placeholder was created for a NO_MATCH name.
    • failureReason replaces identification when a name could not be identified. Send that name alone to POST /v1/lookup.

When the job is accepted it spends one natural-search quota call for each name it may identify, up to its per-article limit. A repeat of the same article that returns the earlier job spends nothing. The job answers 402 and 429 under the same rules as a single lookup.

Example

Set ARTICLE_URL to the article’s address and start a job:

curl -X POST "https://api.aventure.vc/v1/lookup-jobs" \
-H "Authorization: Bearer $AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d '{"sourceUrl":"'"$ARTICLE_URL"'"}'

Read the job with the returned jobId until state is COMPLETED:

curl "https://api.aventure.vc/v1/lookup-jobs/$JOB_ID" \
-H "Authorization: Bearer $AUTH_TOKEN"

Collect the ids from each mention’s match.owner or shell, then queue research on the ones worth updating.