🧬 연구 허브

Skill

batch-author-extractor 코드 참조

저자 역할을 일괄 추출한다. WOS → PubMed → PDF 3단 폴백.

ExtractedAuthorInfo가 없는 모든 PAPER 타입 콘텐츠에 대해 저자 정보를 일괄 추출합니다. WOS(PostgreSQL) → SQLite Content.authors → PDF(GPT Vision) 순서로 시도합니다.

사용법

CLI 스크립트

# 대상 논문 수만 확인 (dry run)
npx tsx scripts/batch-extract-authors.ts --dry-run

# 전체 실행
npx tsx scripts/batch-extract-authors.ts

# 처리할 논문 수 제한
npx tsx scripts/batch-extract-authors.ts --limit 100

# PDF 분석 건너뛰기 (WOS/PubMed 데이터만 사용)
npx tsx scripts/batch-extract-authors.ts --skip-pdf

API 엔드포인트 (개별 논문)

POST /api/papers/{id}/extract-authors
→ 저자 정보 추출 및 저장

GET /api/papers/{id}/extract-authors
→ 저장된 저자 정보 조회

PUT /api/papers/{id}/extract-authors
→ 저자 정보 수동 수정

TypeScript 서비스 직접 호출

import {
  extractAuthorInfo,
  saveExtractedAuthorInfo,
  getExtractedAuthorInfo,
  getWOSDataByDOI,
  parseWOSAuthorData,
  parseContentAuthors
} from "@/lib/author-extraction-service"

// 메인 추출 함수 (WOS → PubMed → PDF 자동 폴백)
const result = await extractAuthorInfo(contentId)

if (result) {
  // DB 저장
  await saveExtractedAuthorInfo(contentId, result)

  console.log(`소스: ${result.source}`)           // "wos" | "pubmed" | "pdf"
  console.log(`제1저자: ${result.firstAuthors}`)
  console.log(`교신저자: ${result.correspondingAuthors}`)
  console.log(`신뢰도: ${result.confidence}`)
}

// WOS 데이터만 직접 조회
const wosData = await getWOSDataByDOI("10.1234/example")
if (wosData) {
  const parsed = parseWOSAuthorData(wosData)
}

3단계 폴백 체인

1단계: WOS (PostgreSQL)

DOI로 Web of Science 데이터를 조회합니다.

DOI → WOS 테이블 순차 조회 (UUS → USN → UYS → USS → UCT → UKR)
→ Author Full Names, Addresses, Reprint Addresses, Email Addresses 파싱

추출 정보:

  • 제1저자 (Author Full Names 첫 번째)
  • 교신저자 (Reprint Addresses에서 "corresponding author" 패턴)
  • 소속 (Addresses 브라켓 파싱)
  • 이메일 (Email Addresses 매칭)
  • 펀딩 (Funding Text 패턴 매칭)

신뢰도: 0.9

2단계: SQLite Content.authors

PubMed에서 가져온 저자 정보를 파싱합니다.

Content.authors (JSON) → authorOrder 정렬
→ 제1저자 (첫 번째), 교신저자 (isCorresponding / 이메일 패턴 / ORCID)

신뢰도: 0.6~0.8 (교신저자 유무에 따라)

3단계: PDF + GPT Vision

PDF가 있으면 항상 실행되어 공동 1저자를 식별합니다.

PDF → Python 스크립트 (extract-authors-from-pdf.py)
→ paper_author_extractor_wos 패키지 또는 GPT-5.2 Vision API
→ 공동1저자 마커 (†, ‡, #), 교신저자 (*) 식별

추출 정보:

  • 공동 제1저자 (dagger/hash 마커)
  • 교신저자 (asterisk 마커)
  • co_first_indicator (사용된 기호)

신뢰도: 0.8~0.85

결과 데이터 구조

interface ExtractedAuthorResult {
  firstAuthors: {
    name: string
    affiliations: string[]
    isCoFirst?: boolean    // 공동 1저자 여부
  }[]
  correspondingAuthors: {
    name: string
    email?: string
    affiliations: string[]
  }[]
  funding?: {
    agency: string
    grantNumber?: string
  }[]
  coFirstIndicator?: string  // "†", "‡", "#" 등
  source: "wos" | "pubmed" | "pdf"
  confidence: number          // 0~1
  pagesAnalyzed?: number
}

환경 변수

변수설명필수
ANALYTICS_DATABASE_URLPostgreSQL WOS DBWOS 소스 사용 시
OPENAI_API_KEYOpenAI API 키PDF 분석 사용 시
OPENAI_MODELVision 모델명 (기본: gpt-5.2)선택

의존성

# Python (PDF 분석용)
pip install PyMuPDF openai

# TypeScript (프로젝트 내장)
# 추가 설치 불필요

WOS 테이블 구조

테이블기관비고
UUS울산대/서울아산병원우선 조회
USN서울대학교
UYS연세대학교
USS다기관 연구
UCT가톨릭대학교
UKRKorea 관련

트러블슈팅

1. WOS 데이터를 못 찾음

원인: DOI가 없거나 WOS에 미등록 해결: PubMed/PDF로 자동 폴백됨

2. PDF 분석 타임아웃

원인: 60초 제한 초과 해결: --skip-pdf 옵션으로 PDF 분석 건너뛰기

3. 공동 1저자 미식별

원인: 논문에 마커가 없거나 비표준 마커 사용 해결: PUT API로 수동 수정

4. 복수 교신저자 누락

원인: WOS Reprint Addresses 파싱 패턴 미매칭 해결: WOS 형식 자동 감지 (세미콜론 또는 쉼표 홀수개)


관련 파일

파일경로설명
추출 서비스src/lib/author-extraction-service.ts3단계 폴백 메인 로직
배치 스크립트scripts/batch-extract-authors.tsCLI 일괄 추출
PDF 추출기scripts/extract-authors-from-pdf.pyPython 래퍼
WOS 패키지~/projects/skills/paper_author_extractor_wos/Python 추출기
PubMed 패키지~/projects/skills/paper_author_extractor_pubmed/Python 추출기
API 라우트src/app/api/papers/[id]/extract-authors/route.ts개별 논문 API

업데이트 이력

  • 2026-02: 스킬 등록
  • 2025-01: WOS 형식 자동 감지 (세미콜론/쉼표), 복수 교신저자 지원