ExtractedAuthorInfo가 없는 모든 PAPER 타입 콘텐츠에 대해 저자 정보를 일괄 추출합니다. WOS(PostgreSQL) → SQLite Content.authors → PDF(GPT Vision) 순서로 시도합니다.
사용법
CLI 스크립트
# 대상 논문 수만 확인 (dry run)
npx tsx scripts/batch-extract-authors.ts --dry-run
# 전체 실행
npx tsx scripts/batch-extract-authors.ts
# 처리할 논문 수 제한
npx tsx scripts/batch-extract-authors.ts --limit 100
# PDF 분석 건너뛰기 (WOS/PubMed 데이터만 사용)
npx tsx scripts/batch-extract-authors.ts --skip-pdfAPI 엔드포인트 (개별 논문)
POST /api/papers/{id}/extract-authors
→ 저자 정보 추출 및 저장
GET /api/papers/{id}/extract-authors
→ 저장된 저자 정보 조회
PUT /api/papers/{id}/extract-authors
→ 저자 정보 수동 수정TypeScript 서비스 직접 호출
import {
extractAuthorInfo,
saveExtractedAuthorInfo,
getExtractedAuthorInfo,
getWOSDataByDOI,
parseWOSAuthorData,
parseContentAuthors
} from "@/lib/author-extraction-service"
// 메인 추출 함수 (WOS → PubMed → PDF 자동 폴백)
const result = await extractAuthorInfo(contentId)
if (result) {
// DB 저장
await saveExtractedAuthorInfo(contentId, result)
console.log(`소스: ${result.source}`) // "wos" | "pubmed" | "pdf"
console.log(`제1저자: ${result.firstAuthors}`)
console.log(`교신저자: ${result.correspondingAuthors}`)
console.log(`신뢰도: ${result.confidence}`)
}
// WOS 데이터만 직접 조회
const wosData = await getWOSDataByDOI("10.1234/example")
if (wosData) {
const parsed = parseWOSAuthorData(wosData)
}3단계 폴백 체인
1단계: WOS (PostgreSQL)
DOI로 Web of Science 데이터를 조회합니다.
DOI → WOS 테이블 순차 조회 (UUS → USN → UYS → USS → UCT → UKR)
→ Author Full Names, Addresses, Reprint Addresses, Email Addresses 파싱추출 정보:
- 제1저자 (Author Full Names 첫 번째)
- 교신저자 (Reprint Addresses에서 "corresponding author" 패턴)
- 소속 (Addresses 브라켓 파싱)
- 이메일 (Email Addresses 매칭)
- 펀딩 (Funding Text 패턴 매칭)
신뢰도: 0.9
2단계: SQLite Content.authors
PubMed에서 가져온 저자 정보를 파싱합니다.
Content.authors (JSON) → authorOrder 정렬
→ 제1저자 (첫 번째), 교신저자 (isCorresponding / 이메일 패턴 / ORCID)신뢰도: 0.6~0.8 (교신저자 유무에 따라)
3단계: PDF + GPT Vision
PDF가 있으면 항상 실행되어 공동 1저자를 식별합니다.
PDF → Python 스크립트 (extract-authors-from-pdf.py)
→ paper_author_extractor_wos 패키지 또는 GPT-5.2 Vision API
→ 공동1저자 마커 (†, ‡, #), 교신저자 (*) 식별추출 정보:
- 공동 제1저자 (dagger/hash 마커)
- 교신저자 (asterisk 마커)
- co_first_indicator (사용된 기호)
신뢰도: 0.8~0.85
결과 데이터 구조
interface ExtractedAuthorResult {
firstAuthors: {
name: string
affiliations: string[]
isCoFirst?: boolean // 공동 1저자 여부
}[]
correspondingAuthors: {
name: string
email?: string
affiliations: string[]
}[]
funding?: {
agency: string
grantNumber?: string
}[]
coFirstIndicator?: string // "†", "‡", "#" 등
source: "wos" | "pubmed" | "pdf"
confidence: number // 0~1
pagesAnalyzed?: number
}환경 변수
| 변수 | 설명 | 필수 |
|---|---|---|
ANALYTICS_DATABASE_URL | PostgreSQL WOS DB | WOS 소스 사용 시 |
OPENAI_API_KEY | OpenAI API 키 | PDF 분석 사용 시 |
OPENAI_MODEL | Vision 모델명 (기본: gpt-5.2) | 선택 |
의존성
# Python (PDF 분석용)
pip install PyMuPDF openai
# TypeScript (프로젝트 내장)
# 추가 설치 불필요WOS 테이블 구조
| 테이블 | 기관 | 비고 |
|---|---|---|
| UUS | 울산대/서울아산병원 | 우선 조회 |
| USN | 서울대학교 | |
| UYS | 연세대학교 | |
| USS | 다기관 연구 | |
| UCT | 가톨릭대학교 | |
| UKR | Korea 관련 |
트러블슈팅
1. WOS 데이터를 못 찾음
원인: DOI가 없거나 WOS에 미등록 해결: PubMed/PDF로 자동 폴백됨
2. PDF 분석 타임아웃
원인: 60초 제한 초과 해결: --skip-pdf 옵션으로 PDF 분석 건너뛰기
3. 공동 1저자 미식별
원인: 논문에 마커가 없거나 비표준 마커 사용 해결: PUT API로 수동 수정
4. 복수 교신저자 누락
원인: WOS Reprint Addresses 파싱 패턴 미매칭 해결: WOS 형식 자동 감지 (세미콜론 또는 쉼표 홀수개)
관련 파일
| 파일 | 경로 | 설명 |
|---|---|---|
| 추출 서비스 | src/lib/author-extraction-service.ts | 3단계 폴백 메인 로직 |
| 배치 스크립트 | scripts/batch-extract-authors.ts | CLI 일괄 추출 |
| PDF 추출기 | scripts/extract-authors-from-pdf.py | Python 래퍼 |
| WOS 패키지 | ~/projects/skills/paper_author_extractor_wos/ | Python 추출기 |
| PubMed 패키지 | ~/projects/skills/paper_author_extractor_pubmed/ | Python 추출기 |
| API 라우트 | src/app/api/papers/[id]/extract-authors/route.ts | 개별 논문 API |
업데이트 이력
- 2026-02: 스킬 등록
- 2025-01: WOS 형식 자동 감지 (세미콜론/쉼표), 복수 교신저자 지원