Add reusable LLM parser profiles with multi-event extract.
Support kind=llm profiles (instruction/schema), optional multi-event posts via #eN URLs, and recover stale running/queued parse jobs after worker crashes. Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
@@ -27,6 +27,9 @@ ADMIN_JWT_SECRET=change-me-jwt-secret
|
|||||||
# DEEPSEEK_BASE_URL=https://api.deepseek.com
|
# DEEPSEEK_BASE_URL=https://api.deepseek.com
|
||||||
# DEEPSEEK_MODEL=deepseek-chat
|
# DEEPSEEK_MODEL=deepseek-chat
|
||||||
|
|
||||||
|
# Recover parse jobs stuck in running/queued after worker crash (seconds, default 900)
|
||||||
|
# STALE_JOB_SECONDS=900
|
||||||
|
|
||||||
# CP adapter workers (set in docker-compose; override locally if needed)
|
# CP adapter workers (set in docker-compose; override locally if needed)
|
||||||
# ENABLED_ADAPTERS=telegram
|
# ENABLED_ADAPTERS=telegram
|
||||||
# WORKER_FAMILIES=telegram
|
# WORKER_FAMILIES=telegram
|
||||||
|
|||||||
@@ -74,7 +74,7 @@ Adapters return dicts matching `IngestEventItem` (`contracts/ingest.py`). Requir
|
|||||||
|
|
||||||
### LLM extract (runtime, not dev)
|
### LLM extract (runtime, not dev)
|
||||||
|
|
||||||
`extract_mode: llm` uses DeepSeek via `workers/llm_extract.py`. Key: `DEEPSEEK_API_KEY` in `.env`. Do not confuse with Cursor dev agents.
|
`extract_mode: llm` uses DeepSeek via `workers/llm_extract.py` (batch). Reusable profiles in admin UI (`kind=heuristic|llm`) flatten into `extract_mode=profile` or `llm`. Key: `DEEPSEEK_API_KEY` in `.env`. Do not confuse with Cursor dev agents.
|
||||||
|
|
||||||
## Common tasks
|
## Common tasks
|
||||||
|
|
||||||
|
|||||||
@@ -76,7 +76,8 @@ docker compose up --build
|
|||||||
| Раздел | Путь | Описание |
|
| Раздел | Путь | Описание |
|
||||||
|--------|------|----------|
|
|--------|------|----------|
|
||||||
| **Карта** | `/` | Интерактивная карта событий (публичный просмотр). CRUD ручных объектов — только для админа (ПКМ). Поддерживает `?eventId=` |
|
| **Карта** | `/` | Интерактивная карта событий (публичный просмотр). CRUD ручных объектов — только для админа (ПКМ). Поддерживает `?eventId=` |
|
||||||
| **Парсеры** | `/parsers` | Адаптеры `telegram` / `crawl4ai` / `viina`, интервал, CRUD; дедуп по `source_url` |
|
| **Парсеры** | `/parsers` | Связка канал + профиль (`heuristic`→`extract_mode=profile`, `llm`→`extract_mode=llm`); дедуп по `source_url` |
|
||||||
|
| **Профили** | `/parser-profiles` | Reusable heuristic (правила) или llm (instruction/schema); целевые поля = Event |
|
||||||
| **События** | `/events` | Фильтрация, пагинация, просмотр деталей, ссылка «На карте» для событий с координатами |
|
| **События** | `/events` | Фильтрация, пагинация, просмотр деталей, ссылка «На карте» для событий с координатами |
|
||||||
| **Аналитика** | `/analytics` | KPI-карточки, график динамики ingest за 30 дней, топ населённых пунктов и регионов |
|
| **Аналитика** | `/analytics` | KPI-карточки, график динамики ingest за 30 дней, топ населённых пунктов и регионов |
|
||||||
| **ПИ** | `/consumers` | CRUD подписчиков distribution API, ротация ключей, тест среза через `/api/v1/events` |
|
| **ПИ** | `/consumers` | CRUD подписчиков distribution API, ротация ключей, тест среза через `/api/v1/events` |
|
||||||
@@ -214,3 +215,8 @@ docker compose down
|
|||||||
```
|
```
|
||||||
|
|
||||||
Данные PostgreSQL сохраняются в volume `pgdata`.
|
Данные PostgreSQL сохраняются в volume `pgdata`.
|
||||||
|
|
||||||
|
## Вход в админку (/login):
|
||||||
|
|
||||||
|
Логин: admin
|
||||||
|
Пароль: change-me
|
||||||
|
|||||||
@@ -98,8 +98,11 @@ class ParserProfile(Base):
|
|||||||
|
|
||||||
id: Mapped[int] = mapped_column(Integer, primary_key=True, index=True)
|
id: Mapped[int] = mapped_column(Integer, primary_key=True, index=True)
|
||||||
name: Mapped[str] = mapped_column(String(255), nullable=False)
|
name: Mapped[str] = mapped_column(String(255), nullable=False)
|
||||||
|
# heuristic = static rules; llm = instruction + extract_schema at CP runtime
|
||||||
|
kind: Mapped[str] = mapped_column(String(50), default="heuristic", nullable=False, index=True)
|
||||||
sample_post: Mapped[str] = mapped_column(Text, default="")
|
sample_post: Mapped[str] = mapped_column(Text, default="")
|
||||||
heuristic_profile: Mapped[dict | None] = mapped_column(JSON, nullable=True)
|
heuristic_profile: Mapped[dict | None] = mapped_column(JSON, nullable=True)
|
||||||
|
llm_profile: Mapped[dict | None] = mapped_column(JSON, nullable=True)
|
||||||
status: Mapped[str] = mapped_column(String(50), default="draft", index=True)
|
status: Mapped[str] = mapped_column(String(50), default="draft", index=True)
|
||||||
created_at: Mapped[datetime] = mapped_column(
|
created_at: Mapped[datetime] = mapped_column(
|
||||||
DateTime(timezone=True),
|
DateTime(timezone=True),
|
||||||
|
|||||||
@@ -87,7 +87,11 @@ def create_parse_job(payload: ParseJobCreate, db: Session = Depends(get_db)):
|
|||||||
raise HTTPException(status_code=404, detail="Profile not found")
|
raise HTTPException(status_code=404, detail="Profile not found")
|
||||||
if not channel:
|
if not channel:
|
||||||
raise HTTPException(status_code=404, detail="Channel not found")
|
raise HTTPException(status_code=404, detail="Channel not found")
|
||||||
if not profile.heuristic_profile:
|
kind = (profile.kind or "heuristic").strip().lower()
|
||||||
|
if kind == "llm":
|
||||||
|
if not profile.llm_profile:
|
||||||
|
raise HTTPException(status_code=400, detail="LLM profile has no llm_profile")
|
||||||
|
elif not profile.heuristic_profile:
|
||||||
raise HTTPException(status_code=400, detail="Profile has no heuristic_profile")
|
raise HTTPException(status_code=400, detail="Profile has no heuristic_profile")
|
||||||
if not channel.is_active:
|
if not channel.is_active:
|
||||||
raise HTTPException(status_code=400, detail="Channel is inactive")
|
raise HTTPException(status_code=400, detail="Channel is inactive")
|
||||||
@@ -146,7 +150,10 @@ def retry_parse_job(job_id: int, db: Session = Depends(get_db)):
|
|||||||
job = db.query(ParseJob).filter(ParseJob.id == job_id).first()
|
job = db.query(ParseJob).filter(ParseJob.id == job_id).first()
|
||||||
if not job:
|
if not job:
|
||||||
raise HTTPException(status_code=404, detail="Job not found")
|
raise HTTPException(status_code=404, detail="Job not found")
|
||||||
if job.status in ("queued", "running"):
|
|
||||||
|
from ..services.job_stale import is_stale_job
|
||||||
|
|
||||||
|
if job.status in ("queued", "running") and not is_stale_job(job):
|
||||||
raise HTTPException(status_code=409, detail="Job is already running or queued")
|
raise HTTPException(status_code=409, detail="Job is already running or queued")
|
||||||
|
|
||||||
job.status = "queued"
|
job.status = "queued"
|
||||||
@@ -218,7 +225,14 @@ def delete_parse_job(job_id: int, db: Session = Depends(get_db)):
|
|||||||
if not job:
|
if not job:
|
||||||
raise HTTPException(status_code=404, detail="Job not found")
|
raise HTTPException(status_code=404, detail="Job not found")
|
||||||
if job.status == "running":
|
if job.status == "running":
|
||||||
raise HTTPException(status_code=409, detail="Cannot delete a running job")
|
from ..services.job_stale import is_stale_job
|
||||||
|
|
||||||
|
if not is_stale_job(job):
|
||||||
|
raise HTTPException(status_code=409, detail="Cannot delete a running job")
|
||||||
|
# Stale running — allow delete after marking failed for audit trail
|
||||||
|
job.status = "failed"
|
||||||
|
job.last_error = "Deleted while stale running"
|
||||||
|
db.commit()
|
||||||
|
|
||||||
db.delete(job)
|
db.delete(job)
|
||||||
db.commit()
|
db.commit()
|
||||||
|
|||||||
@@ -52,7 +52,8 @@ def update_job_status(
|
|||||||
raise HTTPException(status_code=404, detail="Job not found")
|
raise HTTPException(status_code=404, detail="Job not found")
|
||||||
|
|
||||||
job.status = status
|
job.status = status
|
||||||
if status in ("completed", "failed"):
|
# Anchor staleness detection: running/queued start, and terminal finish
|
||||||
|
if status in ("running", "queued", "completed", "failed"):
|
||||||
job.last_run_at = datetime.now(timezone.utc)
|
job.last_run_at = datetime.now(timezone.utc)
|
||||||
job.last_error = error
|
job.last_error = error
|
||||||
db.commit()
|
db.commit()
|
||||||
|
|||||||
@@ -9,6 +9,7 @@ from pydantic import BaseModel, Field
|
|||||||
from sqlalchemy.orm import Session
|
from sqlalchemy.orm import Session
|
||||||
|
|
||||||
from contracts.heuristic_profile import HeuristicProfile
|
from contracts.heuristic_profile import HeuristicProfile
|
||||||
|
from contracts.llm_profile import LlmProfile
|
||||||
|
|
||||||
from ..database import get_db
|
from ..database import get_db
|
||||||
from ..deps import verify_admin
|
from ..deps import verify_admin
|
||||||
@@ -44,6 +45,22 @@ class PreviewRequest(BaseModel):
|
|||||||
|
|
||||||
class PreviewResponse(BaseModel):
|
class PreviewResponse(BaseModel):
|
||||||
fields: dict[str, str]
|
fields: dict[str, str]
|
||||||
|
matched: bool = True
|
||||||
|
missing_required: list[str] = Field(default_factory=list)
|
||||||
|
|
||||||
|
|
||||||
|
class PreviewLlmRequest(BaseModel):
|
||||||
|
sample_post: str = Field(min_length=1)
|
||||||
|
llm_profile: dict[str, Any]
|
||||||
|
|
||||||
|
|
||||||
|
class PreviewLlmResponse(BaseModel):
|
||||||
|
fields: dict[str, str]
|
||||||
|
events: list[dict[str, str]] = Field(default_factory=list)
|
||||||
|
matched: bool = True
|
||||||
|
missing_required: list[str] = Field(default_factory=list)
|
||||||
|
is_event: bool = True
|
||||||
|
matched_count: int = 0
|
||||||
|
|
||||||
|
|
||||||
def _validate_heuristic_profile(raw: dict[str, Any] | None) -> dict[str, Any] | None:
|
def _validate_heuristic_profile(raw: dict[str, Any] | None) -> dict[str, Any] | None:
|
||||||
@@ -55,9 +72,33 @@ def _validate_heuristic_profile(raw: dict[str, Any] | None) -> dict[str, Any] |
|
|||||||
raise HTTPException(status_code=400, detail=f"Invalid heuristic_profile: {exc}") from exc
|
raise HTTPException(status_code=400, detail=f"Invalid heuristic_profile: {exc}") from exc
|
||||||
|
|
||||||
|
|
||||||
def _profile_status(heuristic_profile: dict | None, explicit: str | None = None) -> str:
|
def _validate_llm_profile(raw: dict[str, Any] | None) -> dict[str, Any] | None:
|
||||||
|
if raw is None:
|
||||||
|
return None
|
||||||
|
try:
|
||||||
|
return LlmProfile.model_validate(raw).model_dump()
|
||||||
|
except Exception as exc:
|
||||||
|
raise HTTPException(status_code=400, detail=f"Invalid llm_profile: {exc}") from exc
|
||||||
|
|
||||||
|
|
||||||
|
def _normalize_kind(kind: str | None) -> str:
|
||||||
|
value = (kind or "heuristic").strip().lower()
|
||||||
|
if value not in ("heuristic", "llm"):
|
||||||
|
raise HTTPException(status_code=400, detail="kind must be heuristic or llm")
|
||||||
|
return value
|
||||||
|
|
||||||
|
|
||||||
|
def _profile_status(
|
||||||
|
*,
|
||||||
|
kind: str,
|
||||||
|
heuristic_profile: dict | None,
|
||||||
|
llm_profile: dict | None,
|
||||||
|
explicit: str | None = None,
|
||||||
|
) -> str:
|
||||||
if explicit:
|
if explicit:
|
||||||
return explicit
|
return explicit
|
||||||
|
if kind == "llm":
|
||||||
|
return "ready" if llm_profile else "draft"
|
||||||
return "ready" if heuristic_profile else "draft"
|
return "ready" if heuristic_profile else "draft"
|
||||||
|
|
||||||
|
|
||||||
@@ -66,6 +107,18 @@ def target_fields():
|
|||||||
return {"fields": builder.get_target_fields()}
|
return {"fields": builder.get_target_fields()}
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/llm-defaults")
|
||||||
|
def llm_defaults():
|
||||||
|
from contracts.llm_profile import DEFAULT_EXTRACT_SCHEMA, DEFAULT_INSTRUCTION
|
||||||
|
|
||||||
|
return {
|
||||||
|
"instruction": DEFAULT_INSTRUCTION,
|
||||||
|
"extract_schema": DEFAULT_EXTRACT_SCHEMA,
|
||||||
|
"required_fields": [],
|
||||||
|
"multi_event": False,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
@router.post("/generate", response_model=GenerateResponse)
|
@router.post("/generate", response_model=GenerateResponse)
|
||||||
async def generate_profile(payload: GenerateRequest):
|
async def generate_profile(payload: GenerateRequest):
|
||||||
if not builder.deepseek_enabled():
|
if not builder.deepseek_enabled():
|
||||||
@@ -87,6 +140,9 @@ async def generate_profile(payload: GenerateRequest):
|
|||||||
detail=f"DeepSeek generate failed: {exc}",
|
detail=f"DeepSeek generate failed: {exc}",
|
||||||
) from exc
|
) from exc
|
||||||
dumped = profile.model_dump()
|
dumped = profile.model_dump()
|
||||||
|
if payload.current_profile and isinstance(payload.current_profile.get("required_fields"), list):
|
||||||
|
dumped["required_fields"] = payload.current_profile["required_fields"]
|
||||||
|
dumped = HeuristicProfile.model_validate(dumped).model_dump()
|
||||||
preview = builder.preview_with_profile(payload.sample_post, dumped)
|
preview = builder.preview_with_profile(payload.sample_post, dumped)
|
||||||
return GenerateResponse(
|
return GenerateResponse(
|
||||||
profile=dumped,
|
profile=dumped,
|
||||||
@@ -102,10 +158,40 @@ def preview_profile(payload: PreviewRequest):
|
|||||||
raise HTTPException(status_code=400, detail="heuristic_profile is required")
|
raise HTTPException(status_code=400, detail="heuristic_profile is required")
|
||||||
try:
|
try:
|
||||||
HeuristicProfile.model_validate(raw)
|
HeuristicProfile.model_validate(raw)
|
||||||
fields = builder.preview_with_profile(payload.sample_post, raw)
|
fields, matched, missing = builder.match_preview(payload.sample_post, raw)
|
||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
raise HTTPException(status_code=400, detail=str(exc)) from exc
|
raise HTTPException(status_code=400, detail=str(exc)) from exc
|
||||||
return PreviewResponse(fields=fields)
|
return PreviewResponse(fields=fields, matched=matched, missing_required=missing)
|
||||||
|
|
||||||
|
|
||||||
|
@router.post("/preview-llm", response_model=PreviewLlmResponse)
|
||||||
|
async def preview_llm_profile(payload: PreviewLlmRequest):
|
||||||
|
if not builder.deepseek_enabled():
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=503,
|
||||||
|
detail="DEEPSEEK_API_KEY is not set on ca-api. Add it to .env for LLM preview.",
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
LlmProfile.model_validate(payload.llm_profile)
|
||||||
|
events, fields, matched, missing, is_event, matched_count = await builder.preview_llm_extract(
|
||||||
|
payload.sample_post,
|
||||||
|
payload.llm_profile,
|
||||||
|
)
|
||||||
|
except ValueError as exc:
|
||||||
|
raise HTTPException(status_code=400, detail=str(exc)) from exc
|
||||||
|
except Exception as exc:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=502,
|
||||||
|
detail=f"DeepSeek LLM preview failed: {exc}",
|
||||||
|
) from exc
|
||||||
|
return PreviewLlmResponse(
|
||||||
|
fields=fields,
|
||||||
|
events=events,
|
||||||
|
matched=matched,
|
||||||
|
missing_required=missing,
|
||||||
|
is_event=is_event,
|
||||||
|
matched_count=matched_count,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
@router.get("", response_model=list[ParserProfileRead])
|
@router.get("", response_model=list[ParserProfileRead])
|
||||||
@@ -115,12 +201,26 @@ def list_profiles(db: Session = Depends(get_db)):
|
|||||||
|
|
||||||
@router.post("", response_model=ParserProfileRead, status_code=201)
|
@router.post("", response_model=ParserProfileRead, status_code=201)
|
||||||
def create_profile(payload: ParserProfileCreate, db: Session = Depends(get_db)):
|
def create_profile(payload: ParserProfileCreate, db: Session = Depends(get_db)):
|
||||||
|
kind = _normalize_kind(payload.kind)
|
||||||
heuristic = _validate_heuristic_profile(payload.heuristic_profile)
|
heuristic = _validate_heuristic_profile(payload.heuristic_profile)
|
||||||
|
llm = _validate_llm_profile(payload.llm_profile)
|
||||||
|
if kind == "heuristic" and llm and not heuristic:
|
||||||
|
# ignore stray llm blob when creating heuristic
|
||||||
|
llm = None
|
||||||
|
if kind == "llm" and heuristic and not llm:
|
||||||
|
heuristic = None
|
||||||
profile = ParserProfile(
|
profile = ParserProfile(
|
||||||
name=payload.name.strip(),
|
name=payload.name.strip(),
|
||||||
|
kind=kind,
|
||||||
sample_post=payload.sample_post or "",
|
sample_post=payload.sample_post or "",
|
||||||
heuristic_profile=heuristic,
|
heuristic_profile=heuristic if kind == "heuristic" else None,
|
||||||
status=_profile_status(heuristic, payload.status),
|
llm_profile=llm if kind == "llm" else None,
|
||||||
|
status=_profile_status(
|
||||||
|
kind=kind,
|
||||||
|
heuristic_profile=heuristic if kind == "heuristic" else None,
|
||||||
|
llm_profile=llm if kind == "llm" else None,
|
||||||
|
explicit=payload.status,
|
||||||
|
),
|
||||||
)
|
)
|
||||||
db.add(profile)
|
db.add(profile)
|
||||||
db.commit()
|
db.commit()
|
||||||
@@ -150,12 +250,33 @@ def update_profile(
|
|||||||
profile.name = payload.name.strip()
|
profile.name = payload.name.strip()
|
||||||
if payload.sample_post is not None:
|
if payload.sample_post is not None:
|
||||||
profile.sample_post = payload.sample_post
|
profile.sample_post = payload.sample_post
|
||||||
|
if payload.kind is not None:
|
||||||
|
profile.kind = _normalize_kind(payload.kind)
|
||||||
|
kind = _normalize_kind(profile.kind)
|
||||||
|
|
||||||
if "heuristic_profile" in payload.model_fields_set:
|
if "heuristic_profile" in payload.model_fields_set:
|
||||||
profile.heuristic_profile = _validate_heuristic_profile(payload.heuristic_profile)
|
profile.heuristic_profile = _validate_heuristic_profile(payload.heuristic_profile)
|
||||||
|
if "llm_profile" in payload.model_fields_set:
|
||||||
|
profile.llm_profile = _validate_llm_profile(payload.llm_profile)
|
||||||
|
|
||||||
|
# Keep only the blob matching kind
|
||||||
|
if kind == "heuristic":
|
||||||
|
profile.llm_profile = None
|
||||||
|
else:
|
||||||
|
profile.heuristic_profile = None
|
||||||
|
|
||||||
if payload.status is not None:
|
if payload.status is not None:
|
||||||
profile.status = payload.status
|
profile.status = payload.status
|
||||||
elif "heuristic_profile" in payload.model_fields_set:
|
elif (
|
||||||
profile.status = _profile_status(profile.heuristic_profile)
|
"heuristic_profile" in payload.model_fields_set
|
||||||
|
or "llm_profile" in payload.model_fields_set
|
||||||
|
or payload.kind is not None
|
||||||
|
):
|
||||||
|
profile.status = _profile_status(
|
||||||
|
kind=kind,
|
||||||
|
heuristic_profile=profile.heuristic_profile,
|
||||||
|
llm_profile=profile.llm_profile,
|
||||||
|
)
|
||||||
|
|
||||||
db.commit()
|
db.commit()
|
||||||
db.refresh(profile)
|
db.refresh(profile)
|
||||||
|
|||||||
@@ -161,15 +161,19 @@ class IngestResponse(BaseModel):
|
|||||||
|
|
||||||
class ParserProfileCreate(BaseModel):
|
class ParserProfileCreate(BaseModel):
|
||||||
name: str = Field(min_length=1, max_length=255)
|
name: str = Field(min_length=1, max_length=255)
|
||||||
|
kind: str = Field(default="heuristic", pattern="^(heuristic|llm)$")
|
||||||
sample_post: str = ""
|
sample_post: str = ""
|
||||||
heuristic_profile: dict[str, Any] | None = None
|
heuristic_profile: dict[str, Any] | None = None
|
||||||
|
llm_profile: dict[str, Any] | None = None
|
||||||
status: str | None = None
|
status: str | None = None
|
||||||
|
|
||||||
|
|
||||||
class ParserProfileUpdate(BaseModel):
|
class ParserProfileUpdate(BaseModel):
|
||||||
name: str | None = Field(default=None, min_length=1, max_length=255)
|
name: str | None = Field(default=None, min_length=1, max_length=255)
|
||||||
|
kind: str | None = Field(default=None, pattern="^(heuristic|llm)$")
|
||||||
sample_post: str | None = None
|
sample_post: str | None = None
|
||||||
heuristic_profile: dict[str, Any] | None = None
|
heuristic_profile: dict[str, Any] | None = None
|
||||||
|
llm_profile: dict[str, Any] | None = None
|
||||||
status: str | None = None
|
status: str | None = None
|
||||||
|
|
||||||
|
|
||||||
@@ -178,8 +182,10 @@ class ParserProfileRead(BaseModel):
|
|||||||
|
|
||||||
id: int
|
id: int
|
||||||
name: str
|
name: str
|
||||||
|
kind: str = "heuristic"
|
||||||
sample_post: str
|
sample_post: str
|
||||||
heuristic_profile: dict[str, Any] | None
|
heuristic_profile: dict[str, Any] | None
|
||||||
|
llm_profile: dict[str, Any] | None = None
|
||||||
status: str
|
status: str
|
||||||
created_at: datetime
|
created_at: datetime
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,36 @@
|
|||||||
|
"""Stale parse-job recovery helpers (running/queued left behind after worker crash)."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import os
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
|
||||||
|
from ..models import ParseJob
|
||||||
|
|
||||||
|
# Jobs stuck in running/queued longer than this are considered abandoned.
|
||||||
|
STALE_JOB_SECONDS = int(os.getenv("STALE_JOB_SECONDS", "900"))
|
||||||
|
|
||||||
|
|
||||||
|
def _aware(dt: datetime | None) -> datetime | None:
|
||||||
|
if dt is None:
|
||||||
|
return None
|
||||||
|
if dt.tzinfo is None:
|
||||||
|
return dt.replace(tzinfo=timezone.utc)
|
||||||
|
return dt
|
||||||
|
|
||||||
|
|
||||||
|
def job_anchor_time(job: ParseJob) -> datetime | None:
|
||||||
|
"""Best available timestamp for staleness (prefer last_run_at)."""
|
||||||
|
return _aware(job.last_run_at) or _aware(getattr(job, "created_at", None))
|
||||||
|
|
||||||
|
|
||||||
|
def is_stale_job(job: ParseJob, now: datetime | None = None, *, ttl: int | None = None) -> bool:
|
||||||
|
if job.status not in ("running", "queued"):
|
||||||
|
return False
|
||||||
|
now = now or datetime.now(timezone.utc)
|
||||||
|
anchor = job_anchor_time(job)
|
||||||
|
if anchor is None:
|
||||||
|
# No timestamp — treat long-lived running as stale immediately for recovery
|
||||||
|
return job.status == "running"
|
||||||
|
limit = ttl if ttl is not None else STALE_JOB_SECONDS
|
||||||
|
return (now - anchor).total_seconds() >= limit
|
||||||
@@ -33,6 +33,25 @@ def flatten_pair_config(
|
|||||||
limit: int = 100,
|
limit: int = 100,
|
||||||
) -> dict:
|
) -> dict:
|
||||||
"""Expand Profile + Channel into Redis/CP source_config."""
|
"""Expand Profile + Channel into Redis/CP source_config."""
|
||||||
|
kind = (profile.kind or "heuristic").strip().lower()
|
||||||
|
if kind == "llm":
|
||||||
|
from contracts.llm_profile import LlmProfile
|
||||||
|
|
||||||
|
if not profile.llm_profile:
|
||||||
|
raise ValueError("ParserProfile.llm_profile is empty")
|
||||||
|
llm = LlmProfile.model_validate(profile.llm_profile)
|
||||||
|
cfg = TelegramSourceConfig(
|
||||||
|
channel=channel.channel.strip(),
|
||||||
|
limit=limit,
|
||||||
|
extract_mode="llm",
|
||||||
|
extract_schema=llm.extract_schema,
|
||||||
|
instruction=llm.instruction,
|
||||||
|
required_fields=list(llm.required_fields),
|
||||||
|
multi_event=bool(llm.multi_event),
|
||||||
|
sample_post=profile.sample_post or None,
|
||||||
|
)
|
||||||
|
return cfg.model_dump()
|
||||||
|
|
||||||
if not profile.heuristic_profile:
|
if not profile.heuristic_profile:
|
||||||
raise ValueError("ParserProfile.heuristic_profile is empty")
|
raise ValueError("ParserProfile.heuristic_profile is empty")
|
||||||
cfg = TelegramSourceConfig(
|
cfg = TelegramSourceConfig(
|
||||||
|
|||||||
@@ -23,11 +23,14 @@ def migrate_schema(engine: Engine) -> None:
|
|||||||
if "channel_id" not in columns:
|
if "channel_id" not in columns:
|
||||||
statements.append("ALTER TABLE parse_jobs ADD COLUMN channel_id INTEGER")
|
statements.append("ALTER TABLE parse_jobs ADD COLUMN channel_id INTEGER")
|
||||||
|
|
||||||
# create_all handles new tables; FKs on existing DBs may need indexes
|
if "parser_profiles" in tables:
|
||||||
if "parse_jobs" in tables:
|
columns = {col["name"] for col in inspector.get_columns("parser_profiles")}
|
||||||
# Re-inspect after potential adds is not needed for FK constraints here —
|
if "kind" not in columns:
|
||||||
# create_all + nullable FKs are enough for MVP; optional constraints below.
|
statements.append(
|
||||||
pass
|
"ALTER TABLE parser_profiles ADD COLUMN kind VARCHAR(50) NOT NULL DEFAULT 'heuristic'"
|
||||||
|
)
|
||||||
|
if "llm_profile" not in columns:
|
||||||
|
statements.append("ALTER TABLE parser_profiles ADD COLUMN llm_profile JSON")
|
||||||
|
|
||||||
if not statements:
|
if not statements:
|
||||||
return
|
return
|
||||||
|
|||||||
@@ -14,6 +14,7 @@ from contracts.heuristic_profile import (
|
|||||||
HeuristicProfile,
|
HeuristicProfile,
|
||||||
TARGET_FIELDS,
|
TARGET_FIELDS,
|
||||||
apply_profile,
|
apply_profile,
|
||||||
|
match_profile,
|
||||||
target_field_specs,
|
target_field_specs,
|
||||||
)
|
)
|
||||||
|
|
||||||
@@ -47,6 +48,14 @@ def preview_with_profile(sample_post: str, profile: dict[str, Any] | HeuristicPr
|
|||||||
return apply_profile(sample_post, profile)
|
return apply_profile(sample_post, profile)
|
||||||
|
|
||||||
|
|
||||||
|
def match_preview(
|
||||||
|
sample_post: str,
|
||||||
|
profile: dict[str, Any] | HeuristicProfile,
|
||||||
|
) -> tuple[dict[str, str], bool, list[str]]:
|
||||||
|
matched, fields, missing = match_profile(sample_post, profile)
|
||||||
|
return fields, matched, missing
|
||||||
|
|
||||||
|
|
||||||
def empty_preview_fields(preview: dict[str, str]) -> list[str]:
|
def empty_preview_fields(preview: dict[str, str]) -> list[str]:
|
||||||
return [name for name, value in preview.items() if not (value or "").strip()]
|
return [name for name, value in preview.items() if not (value or "").strip()]
|
||||||
|
|
||||||
@@ -191,3 +200,111 @@ async def generate_profile(
|
|||||||
|
|
||||||
content = data["choices"][0]["message"]["content"]
|
content = data["choices"][0]["message"]["content"]
|
||||||
return _parse_profile_response(content)
|
return _parse_profile_response(content)
|
||||||
|
|
||||||
|
|
||||||
|
async def preview_llm_extract(
|
||||||
|
sample_post: str,
|
||||||
|
llm_profile: dict[str, Any],
|
||||||
|
) -> tuple[list[dict[str, str]], dict[str, str], bool, list[str], bool, int]:
|
||||||
|
"""One-shot DeepSeek extract for CA preview (does not persist).
|
||||||
|
|
||||||
|
Returns (events, fields, matched, missing_required, is_event, matched_count).
|
||||||
|
``fields`` is the first event (or empty) for legacy UI compatibility.
|
||||||
|
"""
|
||||||
|
from contracts.llm_profile import DEFAULT_INSTRUCTION, LlmProfile, match_llm_required
|
||||||
|
|
||||||
|
settings = deepseek_settings()
|
||||||
|
if not settings["api_key"]:
|
||||||
|
raise RuntimeError(
|
||||||
|
"DEEPSEEK_API_KEY is not set on ca-api. Add it to .env for LLM preview."
|
||||||
|
)
|
||||||
|
|
||||||
|
sample = (sample_post or "").strip()
|
||||||
|
if not sample:
|
||||||
|
raise ValueError("sample_post is required")
|
||||||
|
|
||||||
|
profile = LlmProfile.model_validate(llm_profile)
|
||||||
|
schema = profile.extract_schema
|
||||||
|
instr = profile.instruction or DEFAULT_INSTRUCTION
|
||||||
|
schema_lines = "\n".join(f"- {k}: {v}" for k, v in schema.items())
|
||||||
|
multi = bool(profile.multi_event)
|
||||||
|
|
||||||
|
if multi:
|
||||||
|
user_prompt = (
|
||||||
|
f"{instr}\n\n"
|
||||||
|
"If the text describes multiple distinct events (different places, "
|
||||||
|
"coords, or dates), return one object per event in \"events\".\n"
|
||||||
|
f"Fields per event:\n{schema_lines}\n\n"
|
||||||
|
'Return JSON: {"is_event": true|false, "events": [{<field>: <string>}, ...]}\n'
|
||||||
|
"If there is no event, return is_event=false and events=[].\n\n"
|
||||||
|
f"Text:\n{sample[:12000]}"
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
user_prompt = (
|
||||||
|
f"{instr}\n\n"
|
||||||
|
f"Fields to extract:\n{schema_lines}\n\n"
|
||||||
|
'Return JSON: {"is_event": true|false, "fields": {<field>: <string>}}\n\n'
|
||||||
|
f"Text:\n{sample[:12000]}"
|
||||||
|
)
|
||||||
|
|
||||||
|
payload = {
|
||||||
|
"model": settings["model"],
|
||||||
|
"messages": [
|
||||||
|
{
|
||||||
|
"role": "system",
|
||||||
|
"content": (
|
||||||
|
"You extract structured event data for a geoint map. "
|
||||||
|
"Output valid JSON only, no markdown."
|
||||||
|
),
|
||||||
|
},
|
||||||
|
{"role": "user", "content": user_prompt},
|
||||||
|
],
|
||||||
|
"temperature": 0.1,
|
||||||
|
"response_format": {"type": "json_object"},
|
||||||
|
}
|
||||||
|
|
||||||
|
url = f"{settings['base_url']}/chat/completions"
|
||||||
|
async with httpx.AsyncClient(timeout=90.0) as client:
|
||||||
|
response = await client.post(
|
||||||
|
url,
|
||||||
|
headers={
|
||||||
|
"Authorization": f"Bearer {settings['api_key']}",
|
||||||
|
"Content-Type": "application/json",
|
||||||
|
},
|
||||||
|
json=payload,
|
||||||
|
)
|
||||||
|
response.raise_for_status()
|
||||||
|
data = response.json()
|
||||||
|
|
||||||
|
content = data["choices"][0]["message"]["content"]
|
||||||
|
parsed = _extract_json_object(content)
|
||||||
|
is_event = bool(parsed.get("is_event", True))
|
||||||
|
empty_fields = {key: "" for key in schema}
|
||||||
|
|
||||||
|
if not is_event:
|
||||||
|
return [], empty_fields, False, [], False, 0
|
||||||
|
|
||||||
|
events: list[dict[str, str]] = []
|
||||||
|
if multi and isinstance(parsed.get("events"), list):
|
||||||
|
for item in parsed["events"]:
|
||||||
|
if isinstance(item, dict):
|
||||||
|
events.append({key: str(item.get(key) or "").strip() for key in schema})
|
||||||
|
else:
|
||||||
|
fields_raw = parsed.get("fields") if isinstance(parsed.get("fields"), dict) else parsed
|
||||||
|
if not isinstance(fields_raw, dict):
|
||||||
|
fields_raw = {}
|
||||||
|
events.append({key: str(fields_raw.get(key) or "").strip() for key in schema})
|
||||||
|
|
||||||
|
if not events:
|
||||||
|
return [], empty_fields, False, [], False, 0
|
||||||
|
|
||||||
|
matched_events = [ev for ev in events if match_llm_required(ev, list(profile.required_fields))]
|
||||||
|
fields = events[0]
|
||||||
|
missing = [
|
||||||
|
name
|
||||||
|
for name in profile.required_fields
|
||||||
|
if not str(fields.get(name) or "").strip()
|
||||||
|
]
|
||||||
|
matched = match_llm_required(fields, list(profile.required_fields))
|
||||||
|
return events, fields, matched, missing, True, len(matched_events)
|
||||||
|
|
||||||
|
|||||||
@@ -5,6 +5,7 @@ from datetime import datetime, timezone
|
|||||||
|
|
||||||
from ..database import SessionLocal
|
from ..database import SessionLocal
|
||||||
from ..models import ParseJob
|
from ..models import ParseJob
|
||||||
|
from .job_stale import STALE_JOB_SECONDS, is_stale_job
|
||||||
from .jobs import enqueue_parse_job
|
from .jobs import enqueue_parse_job
|
||||||
|
|
||||||
logger = logging.getLogger(__name__)
|
logger = logging.getLogger(__name__)
|
||||||
@@ -13,10 +14,47 @@ TICK_SECONDS = 30
|
|||||||
RECURRING_STATUSES = ("completed", "failed")
|
RECURRING_STATUSES = ("completed", "failed")
|
||||||
|
|
||||||
|
|
||||||
|
def recover_stale_jobs(db, now: datetime) -> None:
|
||||||
|
"""Mark abandoned running/queued jobs as failed; re-queue if still active."""
|
||||||
|
stuck = (
|
||||||
|
db.query(ParseJob)
|
||||||
|
.filter(ParseJob.status.in_(("running", "queued")))
|
||||||
|
.all()
|
||||||
|
)
|
||||||
|
for job in stuck:
|
||||||
|
if not is_stale_job(job, now):
|
||||||
|
continue
|
||||||
|
prev = job.status
|
||||||
|
job.status = "failed"
|
||||||
|
job.last_error = (
|
||||||
|
f"Stale {prev} recovered after {STALE_JOB_SECONDS}s "
|
||||||
|
"(worker likely restarted)"
|
||||||
|
)
|
||||||
|
job.last_run_at = now
|
||||||
|
db.commit()
|
||||||
|
logger.warning("Recovered stale job %s (was %s)", job.id, prev)
|
||||||
|
|
||||||
|
if not job.is_active or job.interval_seconds <= 0:
|
||||||
|
continue
|
||||||
|
job.status = "queued"
|
||||||
|
job.last_error = None
|
||||||
|
db.commit()
|
||||||
|
try:
|
||||||
|
enqueue_parse_job(db, job)
|
||||||
|
logger.info("Re-queued recovered job %s", job.id)
|
||||||
|
except ValueError as exc:
|
||||||
|
job.status = "failed"
|
||||||
|
job.last_error = str(exc)
|
||||||
|
db.commit()
|
||||||
|
logger.warning("Skip re-queue recovered job %s: %s", job.id, exc)
|
||||||
|
|
||||||
|
|
||||||
def run_scheduler_tick() -> None:
|
def run_scheduler_tick() -> None:
|
||||||
db = SessionLocal()
|
db = SessionLocal()
|
||||||
try:
|
try:
|
||||||
now = datetime.now(timezone.utc)
|
now = datetime.now(timezone.utc)
|
||||||
|
recover_stale_jobs(db, now)
|
||||||
|
|
||||||
jobs = (
|
jobs = (
|
||||||
db.query(ParseJob)
|
db.query(ParseJob)
|
||||||
.filter(
|
.filter(
|
||||||
@@ -61,5 +99,9 @@ def start_scheduler() -> threading.Event:
|
|||||||
stop_event = threading.Event()
|
stop_event = threading.Event()
|
||||||
thread = threading.Thread(target=_scheduler_loop, args=(stop_event,), daemon=True)
|
thread = threading.Thread(target=_scheduler_loop, args=(stop_event,), daemon=True)
|
||||||
thread.start()
|
thread.start()
|
||||||
logger.info("Parse job scheduler started (tick every %ss)", TICK_SECONDS)
|
logger.info(
|
||||||
|
"Parse job scheduler started (tick every %ss, stale after %ss)",
|
||||||
|
TICK_SECONDS,
|
||||||
|
STALE_JOB_SECONDS,
|
||||||
|
)
|
||||||
return stop_event
|
return stop_event
|
||||||
|
|||||||
@@ -1,6 +1,7 @@
|
|||||||
import { ADMIN_BASE, request } from "./client";
|
import { ADMIN_BASE, request } from "./client";
|
||||||
import type {
|
import type {
|
||||||
HeuristicProfile,
|
HeuristicProfile,
|
||||||
|
LlmProfile,
|
||||||
ParseChannel,
|
ParseChannel,
|
||||||
ParseChannelCreate,
|
ParseChannelCreate,
|
||||||
ParseChannelUpdate,
|
ParseChannelUpdate,
|
||||||
@@ -10,7 +11,7 @@ import type {
|
|||||||
TargetField,
|
TargetField,
|
||||||
} from "../types/admin";
|
} from "../types/admin";
|
||||||
|
|
||||||
export type { HeuristicProfile, TargetField } from "../types/admin";
|
export type { HeuristicProfile, LlmProfile, TargetField } from "../types/admin";
|
||||||
|
|
||||||
export function fetchChannels(): Promise<ParseChannel[]> {
|
export function fetchChannels(): Promise<ParseChannel[]> {
|
||||||
return request<ParseChannel[]>("/parse-channels", undefined, ADMIN_BASE);
|
return request<ParseChannel[]>("/parse-channels", undefined, ADMIN_BASE);
|
||||||
@@ -64,6 +65,10 @@ export function fetchTargetFields(): Promise<{ fields: TargetField[] }> {
|
|||||||
return request<{ fields: TargetField[] }>("/parser-profiles/target-fields", undefined, ADMIN_BASE);
|
return request<{ fields: TargetField[] }>("/parser-profiles/target-fields", undefined, ADMIN_BASE);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
export function fetchLlmDefaults(): Promise<LlmProfile> {
|
||||||
|
return request<LlmProfile>("/parser-profiles/llm-defaults", undefined, ADMIN_BASE);
|
||||||
|
}
|
||||||
|
|
||||||
export function generateParserProfile(payload: {
|
export function generateParserProfile(payload: {
|
||||||
sample_post: string;
|
sample_post: string;
|
||||||
hint?: string;
|
hint?: string;
|
||||||
@@ -94,10 +99,35 @@ export function generateParserProfile(payload: {
|
|||||||
export function previewParserProfile(
|
export function previewParserProfile(
|
||||||
sample_post: string,
|
sample_post: string,
|
||||||
heuristic_profile: HeuristicProfile | Record<string, unknown>,
|
heuristic_profile: HeuristicProfile | Record<string, unknown>,
|
||||||
): Promise<{ fields: Record<string, string> }> {
|
): Promise<{ fields: Record<string, string>; matched: boolean; missing_required: string[] }> {
|
||||||
return request<{ fields: Record<string, string> }>(
|
return request<{ fields: Record<string, string>; matched: boolean; missing_required: string[] }>(
|
||||||
"/parser-profiles/preview",
|
"/parser-profiles/preview",
|
||||||
{ method: "POST", body: JSON.stringify({ sample_post, heuristic_profile }) },
|
{ method: "POST", body: JSON.stringify({ sample_post, heuristic_profile }) },
|
||||||
ADMIN_BASE,
|
ADMIN_BASE,
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
export function previewLlmProfile(
|
||||||
|
sample_post: string,
|
||||||
|
llm_profile: LlmProfile | Record<string, unknown>,
|
||||||
|
): Promise<{
|
||||||
|
fields: Record<string, string>;
|
||||||
|
events: Record<string, string>[];
|
||||||
|
matched: boolean;
|
||||||
|
missing_required: string[];
|
||||||
|
is_event: boolean;
|
||||||
|
matched_count: number;
|
||||||
|
}> {
|
||||||
|
return request<{
|
||||||
|
fields: Record<string, string>;
|
||||||
|
events: Record<string, string>[];
|
||||||
|
matched: boolean;
|
||||||
|
missing_required: string[];
|
||||||
|
is_event: boolean;
|
||||||
|
matched_count: number;
|
||||||
|
}>(
|
||||||
|
"/parser-profiles/preview-llm",
|
||||||
|
{ method: "POST", body: JSON.stringify({ sample_post, llm_profile }) },
|
||||||
|
ADMIN_BASE,
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|||||||
@@ -36,26 +36,52 @@ export interface ParseJobUpdate {
|
|||||||
is_active?: boolean;
|
is_active?: boolean;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
export type TargetField = {
|
||||||
|
name: string;
|
||||||
|
type: string;
|
||||||
|
description: string;
|
||||||
|
};
|
||||||
|
|
||||||
|
export type HeuristicProfile = {
|
||||||
|
version: number;
|
||||||
|
fields: Record<string, Record<string, unknown>>;
|
||||||
|
required_fields?: string[];
|
||||||
|
notes?: string;
|
||||||
|
};
|
||||||
|
|
||||||
|
export type LlmProfile = {
|
||||||
|
instruction?: string | null;
|
||||||
|
extract_schema: Record<string, string>;
|
||||||
|
required_fields?: string[];
|
||||||
|
multi_event?: boolean;
|
||||||
|
};
|
||||||
|
|
||||||
export interface ParserProfile {
|
export interface ParserProfile {
|
||||||
id: number;
|
id: number;
|
||||||
name: string;
|
name: string;
|
||||||
|
kind: "heuristic" | "llm";
|
||||||
sample_post: string;
|
sample_post: string;
|
||||||
heuristic_profile: Record<string, unknown> | null;
|
heuristic_profile: Record<string, unknown> | null;
|
||||||
|
llm_profile: LlmProfile | Record<string, unknown> | null;
|
||||||
status: string;
|
status: string;
|
||||||
created_at: string;
|
created_at: string;
|
||||||
}
|
}
|
||||||
|
|
||||||
export interface ParserProfileCreate {
|
export interface ParserProfileCreate {
|
||||||
name: string;
|
name: string;
|
||||||
|
kind?: "heuristic" | "llm";
|
||||||
sample_post?: string;
|
sample_post?: string;
|
||||||
heuristic_profile?: Record<string, unknown> | null;
|
heuristic_profile?: Record<string, unknown> | null;
|
||||||
|
llm_profile?: LlmProfile | Record<string, unknown> | null;
|
||||||
status?: string;
|
status?: string;
|
||||||
}
|
}
|
||||||
|
|
||||||
export interface ParserProfileUpdate {
|
export interface ParserProfileUpdate {
|
||||||
name?: string;
|
name?: string;
|
||||||
|
kind?: "heuristic" | "llm";
|
||||||
sample_post?: string;
|
sample_post?: string;
|
||||||
heuristic_profile?: Record<string, unknown> | null;
|
heuristic_profile?: Record<string, unknown> | null;
|
||||||
|
llm_profile?: LlmProfile | Record<string, unknown> | null;
|
||||||
status?: string;
|
status?: string;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -82,18 +108,6 @@ export interface ParseChannelUpdate {
|
|||||||
is_active?: boolean;
|
is_active?: boolean;
|
||||||
}
|
}
|
||||||
|
|
||||||
export type TargetField = {
|
|
||||||
name: string;
|
|
||||||
type: string;
|
|
||||||
description: string;
|
|
||||||
};
|
|
||||||
|
|
||||||
export type HeuristicProfile = {
|
|
||||||
version: number;
|
|
||||||
fields: Record<string, Record<string, unknown>>;
|
|
||||||
notes?: string;
|
|
||||||
};
|
|
||||||
|
|
||||||
export interface EventRecord {
|
export interface EventRecord {
|
||||||
id: number;
|
id: number;
|
||||||
source_type: string;
|
source_type: string;
|
||||||
|
|||||||
@@ -1,14 +1,17 @@
|
|||||||
<script setup lang="ts">
|
<script setup lang="ts">
|
||||||
import { onMounted, ref } from "vue";
|
import { onMounted, ref, watch } from "vue";
|
||||||
import {
|
import {
|
||||||
createProfile,
|
createProfile,
|
||||||
deleteProfile,
|
deleteProfile,
|
||||||
|
fetchLlmDefaults,
|
||||||
fetchProfiles,
|
fetchProfiles,
|
||||||
fetchTargetFields,
|
fetchTargetFields,
|
||||||
generateParserProfile,
|
generateParserProfile,
|
||||||
|
previewLlmProfile,
|
||||||
previewParserProfile,
|
previewParserProfile,
|
||||||
updateProfile,
|
updateProfile,
|
||||||
type HeuristicProfile,
|
type HeuristicProfile,
|
||||||
|
type LlmProfile,
|
||||||
type TargetField,
|
type TargetField,
|
||||||
} from "../api/entities";
|
} from "../api/entities";
|
||||||
import type { ParserProfile } from "../types/admin";
|
import type { ParserProfile } from "../types/admin";
|
||||||
@@ -19,18 +22,29 @@ const error = ref("");
|
|||||||
const success = ref("");
|
const success = ref("");
|
||||||
const submitting = ref(false);
|
const submitting = ref(false);
|
||||||
|
|
||||||
|
const profileKind = ref<"heuristic" | "llm">("heuristic");
|
||||||
const samplePost = ref("");
|
const samplePost = ref("");
|
||||||
const generateHint = ref("");
|
const generateHint = ref("");
|
||||||
const profileName = ref("");
|
const profileName = ref("");
|
||||||
const targetFields = ref<TargetField[]>([]);
|
const targetFields = ref<TargetField[]>([]);
|
||||||
const profileJson = ref("");
|
const profileJson = ref("");
|
||||||
|
const llmInstruction = ref("");
|
||||||
|
const llmSchemaJson = ref("");
|
||||||
|
const llmMultiEvent = ref(false);
|
||||||
const previewFields = ref<Record<string, string> | null>(null);
|
const previewFields = ref<Record<string, string> | null>(null);
|
||||||
|
const previewEvents = ref<Record<string, string>[]>([]);
|
||||||
const emptyFields = ref<string[]>([]);
|
const emptyFields = ref<string[]>([]);
|
||||||
|
const requiredFields = ref<string[]>([]);
|
||||||
|
const previewMatched = ref(true);
|
||||||
|
const previewMissing = ref<string[]>([]);
|
||||||
|
const previewIsEvent = ref(true);
|
||||||
|
const previewMatchedCount = ref(0);
|
||||||
const generating = ref(false);
|
const generating = ref(false);
|
||||||
const previewing = ref(false);
|
const previewing = ref(false);
|
||||||
const loadingFields = ref(true);
|
const loadingFields = ref(true);
|
||||||
const editingId = ref<number | null>(null);
|
const editingId = ref<number | null>(null);
|
||||||
const profileStatus = ref<"draft" | "ready">("draft");
|
const profileStatus = ref<"draft" | "ready">("draft");
|
||||||
|
const llmDefaults = ref<LlmProfile | null>(null);
|
||||||
|
|
||||||
async function loadProfiles() {
|
async function loadProfiles() {
|
||||||
loading.value = true;
|
loading.value = true;
|
||||||
@@ -56,7 +70,28 @@ async function loadTargetFields() {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
function syncProfileFromJson(): HeuristicProfile | null {
|
async function loadLlmDefaults() {
|
||||||
|
try {
|
||||||
|
llmDefaults.value = await fetchLlmDefaults();
|
||||||
|
} catch {
|
||||||
|
llmDefaults.value = null;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function applyLlmDefaults() {
|
||||||
|
const d = llmDefaults.value;
|
||||||
|
if (!d) return;
|
||||||
|
llmInstruction.value = d.instruction || "";
|
||||||
|
llmSchemaJson.value = JSON.stringify(d.extract_schema || {}, null, 2);
|
||||||
|
if (!requiredFields.value.length) {
|
||||||
|
requiredFields.value = Array.isArray(d.required_fields) ? [...d.required_fields] : [];
|
||||||
|
}
|
||||||
|
if (typeof d.multi_event === "boolean") {
|
||||||
|
llmMultiEvent.value = d.multi_event;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function syncHeuristicFromJson(): HeuristicProfile | null {
|
||||||
if (!profileJson.value.trim()) return null;
|
if (!profileJson.value.trim()) return null;
|
||||||
try {
|
try {
|
||||||
return JSON.parse(profileJson.value) as HeuristicProfile;
|
return JSON.parse(profileJson.value) as HeuristicProfile;
|
||||||
@@ -66,7 +101,88 @@ function syncProfileFromJson(): HeuristicProfile | null {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
function syncLlmFromForm(): LlmProfile | null {
|
||||||
|
let schema: Record<string, string>;
|
||||||
|
try {
|
||||||
|
schema = JSON.parse(llmSchemaJson.value || "{}") as Record<string, string>;
|
||||||
|
} catch {
|
||||||
|
error.value = "Некорректный JSON extract_schema";
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
if (!schema || typeof schema !== "object" || !Object.keys(schema).length) {
|
||||||
|
error.value = "extract_schema не должен быть пустым";
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
return {
|
||||||
|
instruction: llmInstruction.value.trim() || null,
|
||||||
|
extract_schema: schema,
|
||||||
|
required_fields: [...requiredFields.value],
|
||||||
|
multi_event: llmMultiEvent.value,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
const DEFAULT_REQUIRED = ["event_date", "coords"];
|
||||||
|
|
||||||
|
function uniqueFields(names: string[]): string[] {
|
||||||
|
return names.filter((name, idx) => names.indexOf(name) === idx);
|
||||||
|
}
|
||||||
|
|
||||||
|
function writeRequiredToHeuristicJson(names: string[]) {
|
||||||
|
const current = syncHeuristicFromJson();
|
||||||
|
if (!current) return;
|
||||||
|
current.required_fields = names;
|
||||||
|
profileJson.value = JSON.stringify(current, null, 2);
|
||||||
|
}
|
||||||
|
|
||||||
|
function requiredFromHeuristic(
|
||||||
|
profile: HeuristicProfile,
|
||||||
|
preview: Record<string, string> | null,
|
||||||
|
): string[] {
|
||||||
|
if (Array.isArray(profile.required_fields)) {
|
||||||
|
return uniqueFields(profile.required_fields);
|
||||||
|
}
|
||||||
|
if (!preview) return [];
|
||||||
|
return DEFAULT_REQUIRED.filter((name) => !!(preview[name] || "").trim());
|
||||||
|
}
|
||||||
|
|
||||||
|
function isRequired(name: string): boolean {
|
||||||
|
return requiredFields.value.includes(name);
|
||||||
|
}
|
||||||
|
|
||||||
|
async function toggleRequired(name: string) {
|
||||||
|
const next = isRequired(name)
|
||||||
|
? requiredFields.value.filter((n) => n !== name)
|
||||||
|
: [...requiredFields.value, name];
|
||||||
|
requiredFields.value = next;
|
||||||
|
if (profileKind.value === "heuristic") {
|
||||||
|
writeRequiredToHeuristicJson(next);
|
||||||
|
}
|
||||||
|
if (previewFields.value && samplePost.value.trim()) {
|
||||||
|
await handlePreview();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
watch(profileKind, (kind, prev) => {
|
||||||
|
if (kind === prev) return;
|
||||||
|
previewFields.value = null;
|
||||||
|
previewEvents.value = [];
|
||||||
|
emptyFields.value = [];
|
||||||
|
previewMatched.value = true;
|
||||||
|
previewMissing.value = [];
|
||||||
|
previewIsEvent.value = true;
|
||||||
|
previewMatchedCount.value = 0;
|
||||||
|
success.value = "";
|
||||||
|
if (kind === "llm" && !llmSchemaJson.value.trim()) {
|
||||||
|
applyLlmDefaults();
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
async function handleGenerate() {
|
async function handleGenerate() {
|
||||||
|
if (profileKind.value === "llm") {
|
||||||
|
applyLlmDefaults();
|
||||||
|
success.value = "Подставлены дефолты LLM schema/instruction. Отредактируйте и Preview.";
|
||||||
|
return;
|
||||||
|
}
|
||||||
if (!samplePost.value.trim()) {
|
if (!samplePost.value.trim()) {
|
||||||
error.value = "Вставьте образец поста";
|
error.value = "Вставьте образец поста";
|
||||||
return;
|
return;
|
||||||
@@ -79,7 +195,7 @@ async function handleGenerate() {
|
|||||||
try {
|
try {
|
||||||
let current: HeuristicProfile | null = null;
|
let current: HeuristicProfile | null = null;
|
||||||
if (profileJson.value.trim()) {
|
if (profileJson.value.trim()) {
|
||||||
current = syncProfileFromJson();
|
current = syncHeuristicFromJson();
|
||||||
if (!current) return;
|
if (!current) return;
|
||||||
}
|
}
|
||||||
const data = await generateParserProfile({
|
const data = await generateParserProfile({
|
||||||
@@ -87,9 +203,18 @@ async function handleGenerate() {
|
|||||||
hint: generateHint.value.trim() || undefined,
|
hint: generateHint.value.trim() || undefined,
|
||||||
current_profile: current,
|
current_profile: current,
|
||||||
});
|
});
|
||||||
|
const required =
|
||||||
|
current && Array.isArray(current.required_fields)
|
||||||
|
? uniqueFields(current.required_fields)
|
||||||
|
: DEFAULT_REQUIRED.filter((name) => !!(data.preview[name] || "").trim());
|
||||||
|
data.profile.required_fields = required;
|
||||||
profileJson.value = JSON.stringify(data.profile, null, 2);
|
profileJson.value = JSON.stringify(data.profile, null, 2);
|
||||||
|
requiredFields.value = required;
|
||||||
previewFields.value = data.preview;
|
previewFields.value = data.preview;
|
||||||
emptyFields.value = data.empty_fields || [];
|
emptyFields.value = data.empty_fields || [];
|
||||||
|
const previewData = await previewParserProfile(samplePost.value, data.profile);
|
||||||
|
previewMatched.value = previewData.matched;
|
||||||
|
previewMissing.value = previewData.missing_required || [];
|
||||||
if (emptyFields.value.length) {
|
if (emptyFields.value.length) {
|
||||||
success.value =
|
success.value =
|
||||||
`Профиль обновлён, но пустые поля: ${emptyFields.value.join(", ")}. ` +
|
`Профиль обновлён, но пустые поля: ${emptyFields.value.join(", ")}. ` +
|
||||||
@@ -109,11 +234,6 @@ async function handleGenerate() {
|
|||||||
}
|
}
|
||||||
|
|
||||||
async function handlePreview() {
|
async function handlePreview() {
|
||||||
const current = syncProfileFromJson();
|
|
||||||
if (!current) {
|
|
||||||
if (!error.value) error.value = "Сначала сгенерируйте или вставьте профиль";
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
if (!samplePost.value.trim()) {
|
if (!samplePost.value.trim()) {
|
||||||
error.value = "Нужен образец поста для preview";
|
error.value = "Нужен образец поста для preview";
|
||||||
return;
|
return;
|
||||||
@@ -121,11 +241,53 @@ async function handlePreview() {
|
|||||||
previewing.value = true;
|
previewing.value = true;
|
||||||
error.value = "";
|
error.value = "";
|
||||||
try {
|
try {
|
||||||
const data = await previewParserProfile(samplePost.value, current);
|
if (profileKind.value === "llm") {
|
||||||
previewFields.value = data.fields;
|
const llm = syncLlmFromForm();
|
||||||
emptyFields.value = Object.entries(data.fields)
|
if (!llm) return;
|
||||||
.filter(([, v]) => !v)
|
const data = await previewLlmProfile(samplePost.value, llm);
|
||||||
.map(([k]) => k);
|
const rows =
|
||||||
|
Array.isArray(data.events) && data.events.length
|
||||||
|
? data.events
|
||||||
|
: data.fields
|
||||||
|
? [data.fields]
|
||||||
|
: [];
|
||||||
|
previewEvents.value = rows;
|
||||||
|
previewFields.value = rows[0] || data.fields || null;
|
||||||
|
emptyFields.value = previewFields.value
|
||||||
|
? Object.entries(previewFields.value)
|
||||||
|
.filter(([, v]) => !v)
|
||||||
|
.map(([k]) => k)
|
||||||
|
: [];
|
||||||
|
previewMatched.value = data.matched;
|
||||||
|
previewMissing.value = data.missing_required || [];
|
||||||
|
previewIsEvent.value = data.is_event;
|
||||||
|
previewMatchedCount.value = data.matched_count ?? rows.length;
|
||||||
|
if (!data.is_event) {
|
||||||
|
success.value = "Модель пометила текст как не-событие (is_event=false) — runtime пропустит.";
|
||||||
|
} else if (rows.length > 1) {
|
||||||
|
success.value = `Preview: ${rows.length} событий из поста (пройдут фильтр: ${previewMatchedCount.value}).`;
|
||||||
|
}
|
||||||
|
} else {
|
||||||
|
const current = syncHeuristicFromJson();
|
||||||
|
if (!current) {
|
||||||
|
if (!error.value) error.value = "Сначала сгенерируйте или вставьте профиль";
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
const data = await previewParserProfile(samplePost.value, current);
|
||||||
|
previewFields.value = data.fields;
|
||||||
|
emptyFields.value = Object.entries(data.fields)
|
||||||
|
.filter(([, v]) => !v)
|
||||||
|
.map(([k]) => k);
|
||||||
|
requiredFields.value = requiredFromHeuristic(current, data.fields);
|
||||||
|
if (!Array.isArray(current.required_fields)) {
|
||||||
|
writeRequiredToHeuristicJson(requiredFields.value);
|
||||||
|
}
|
||||||
|
previewMatched.value = data.matched;
|
||||||
|
previewMissing.value = data.missing_required || [];
|
||||||
|
previewIsEvent.value = true;
|
||||||
|
previewEvents.value = data.fields ? [data.fields] : [];
|
||||||
|
previewMatchedCount.value = data.matched ? 1 : 0;
|
||||||
|
}
|
||||||
} catch (err) {
|
} catch (err) {
|
||||||
error.value = err instanceof Error ? err.message : "Ошибка preview";
|
error.value = err instanceof Error ? err.message : "Ошибка preview";
|
||||||
} finally {
|
} finally {
|
||||||
@@ -135,29 +297,72 @@ async function handlePreview() {
|
|||||||
|
|
||||||
function resetEditor() {
|
function resetEditor() {
|
||||||
editingId.value = null;
|
editingId.value = null;
|
||||||
|
profileKind.value = "heuristic";
|
||||||
profileName.value = "";
|
profileName.value = "";
|
||||||
samplePost.value = "";
|
samplePost.value = "";
|
||||||
generateHint.value = "";
|
generateHint.value = "";
|
||||||
profileJson.value = "";
|
profileJson.value = "";
|
||||||
|
llmInstruction.value = "";
|
||||||
|
llmSchemaJson.value = "";
|
||||||
|
llmMultiEvent.value = false;
|
||||||
profileStatus.value = "draft";
|
profileStatus.value = "draft";
|
||||||
previewFields.value = null;
|
previewFields.value = null;
|
||||||
|
previewEvents.value = [];
|
||||||
emptyFields.value = [];
|
emptyFields.value = [];
|
||||||
|
requiredFields.value = [];
|
||||||
|
previewMatched.value = true;
|
||||||
|
previewMissing.value = [];
|
||||||
|
previewIsEvent.value = true;
|
||||||
|
previewMatchedCount.value = 0;
|
||||||
success.value = "";
|
success.value = "";
|
||||||
}
|
}
|
||||||
|
|
||||||
function openEdit(p: ParserProfile) {
|
function openEdit(p: ParserProfile) {
|
||||||
editingId.value = p.id;
|
editingId.value = p.id;
|
||||||
|
profileKind.value = p.kind === "llm" ? "llm" : "heuristic";
|
||||||
profileName.value = p.name;
|
profileName.value = p.name;
|
||||||
samplePost.value = p.sample_post || "";
|
samplePost.value = p.sample_post || "";
|
||||||
generateHint.value = "";
|
generateHint.value = "";
|
||||||
profileJson.value = p.heuristic_profile
|
|
||||||
? JSON.stringify(p.heuristic_profile, null, 2)
|
|
||||||
: "";
|
|
||||||
profileStatus.value = p.status === "ready" ? "ready" : "draft";
|
profileStatus.value = p.status === "ready" ? "ready" : "draft";
|
||||||
previewFields.value = null;
|
previewFields.value = null;
|
||||||
|
previewEvents.value = [];
|
||||||
emptyFields.value = [];
|
emptyFields.value = [];
|
||||||
|
previewMatched.value = true;
|
||||||
|
previewMissing.value = [];
|
||||||
|
previewIsEvent.value = true;
|
||||||
|
previewMatchedCount.value = 0;
|
||||||
success.value = "";
|
success.value = "";
|
||||||
error.value = "";
|
error.value = "";
|
||||||
|
|
||||||
|
if (p.kind === "llm") {
|
||||||
|
profileJson.value = "";
|
||||||
|
const lp = (p.llm_profile || {}) as LlmProfile;
|
||||||
|
llmInstruction.value = lp.instruction || "";
|
||||||
|
llmSchemaJson.value = JSON.stringify(lp.extract_schema || {}, null, 2);
|
||||||
|
llmMultiEvent.value = !!lp.multi_event;
|
||||||
|
requiredFields.value = Array.isArray(lp.required_fields)
|
||||||
|
? uniqueFields(lp.required_fields)
|
||||||
|
: [];
|
||||||
|
if (!llmSchemaJson.value.trim() || llmSchemaJson.value === "{}") {
|
||||||
|
applyLlmDefaults();
|
||||||
|
}
|
||||||
|
} else {
|
||||||
|
llmInstruction.value = "";
|
||||||
|
llmSchemaJson.value = "";
|
||||||
|
llmMultiEvent.value = false;
|
||||||
|
profileJson.value = p.heuristic_profile
|
||||||
|
? JSON.stringify(p.heuristic_profile, null, 2)
|
||||||
|
: "";
|
||||||
|
const hp = p.heuristic_profile;
|
||||||
|
if (hp && Array.isArray(hp.required_fields)) {
|
||||||
|
requiredFields.value = uniqueFields(hp.required_fields as string[]);
|
||||||
|
} else if (hp) {
|
||||||
|
requiredFields.value = [...DEFAULT_REQUIRED];
|
||||||
|
writeRequiredToHeuristicJson(requiredFields.value);
|
||||||
|
} else {
|
||||||
|
requiredFields.value = [];
|
||||||
|
}
|
||||||
|
}
|
||||||
window.scrollTo({ top: 0, behavior: "smooth" });
|
window.scrollTo({ top: 0, behavior: "smooth" });
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -166,33 +371,60 @@ async function handleSave() {
|
|||||||
error.value = "Укажите название профиля";
|
error.value = "Укажите название профиля";
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
const current = syncProfileFromJson();
|
|
||||||
if (!current) {
|
|
||||||
if (!error.value) error.value = "Нужен JSON профиля (Generate или вручную)";
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
submitting.value = true;
|
submitting.value = true;
|
||||||
error.value = "";
|
error.value = "";
|
||||||
success.value = "";
|
success.value = "";
|
||||||
try {
|
try {
|
||||||
if (editingId.value != null) {
|
if (profileKind.value === "llm") {
|
||||||
await updateProfile(editingId.value, {
|
const llm = syncLlmFromForm();
|
||||||
|
if (!llm) return;
|
||||||
|
const payload = {
|
||||||
name: profileName.value.trim(),
|
name: profileName.value.trim(),
|
||||||
|
kind: "llm" as const,
|
||||||
sample_post: samplePost.value,
|
sample_post: samplePost.value,
|
||||||
heuristic_profile: current,
|
llm_profile: llm,
|
||||||
|
heuristic_profile: null,
|
||||||
status: profileStatus.value,
|
status: profileStatus.value,
|
||||||
});
|
};
|
||||||
success.value = `Профиль #${editingId.value} обновлён`;
|
if (editingId.value != null) {
|
||||||
|
await updateProfile(editingId.value, payload);
|
||||||
|
success.value = `LLM-профиль #${editingId.value} обновлён`;
|
||||||
|
} else {
|
||||||
|
const created = await createProfile({
|
||||||
|
...payload,
|
||||||
|
status: profileStatus.value === "draft" ? "draft" : "ready",
|
||||||
|
});
|
||||||
|
success.value = `LLM-профиль #${created.id} сохранён`;
|
||||||
|
editingId.value = created.id;
|
||||||
|
profileStatus.value = created.status === "ready" ? "ready" : "draft";
|
||||||
|
}
|
||||||
} else {
|
} else {
|
||||||
const created = await createProfile({
|
writeRequiredToHeuristicJson(requiredFields.value);
|
||||||
|
const current = syncHeuristicFromJson();
|
||||||
|
if (!current) {
|
||||||
|
if (!error.value) error.value = "Нужен JSON профиля (Generate или вручную)";
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
const payload = {
|
||||||
name: profileName.value.trim(),
|
name: profileName.value.trim(),
|
||||||
|
kind: "heuristic" as const,
|
||||||
sample_post: samplePost.value,
|
sample_post: samplePost.value,
|
||||||
heuristic_profile: current,
|
heuristic_profile: current,
|
||||||
status: profileStatus.value === "draft" ? "draft" : "ready",
|
llm_profile: null,
|
||||||
});
|
status: profileStatus.value,
|
||||||
success.value = `Профиль #${created.id} сохранён`;
|
};
|
||||||
editingId.value = created.id;
|
if (editingId.value != null) {
|
||||||
profileStatus.value = created.status === "ready" ? "ready" : "draft";
|
await updateProfile(editingId.value, payload);
|
||||||
|
success.value = `Профиль #${editingId.value} обновлён`;
|
||||||
|
} else {
|
||||||
|
const created = await createProfile({
|
||||||
|
...payload,
|
||||||
|
status: profileStatus.value === "draft" ? "draft" : "ready",
|
||||||
|
});
|
||||||
|
success.value = `Профиль #${created.id} сохранён`;
|
||||||
|
editingId.value = created.id;
|
||||||
|
profileStatus.value = created.status === "ready" ? "ready" : "draft";
|
||||||
|
}
|
||||||
}
|
}
|
||||||
await loadProfiles();
|
await loadProfiles();
|
||||||
} catch (err) {
|
} catch (err) {
|
||||||
@@ -218,9 +450,18 @@ function formatDate(value: string): string {
|
|||||||
return new Date(value).toLocaleString("ru-RU");
|
return new Date(value).toLocaleString("ru-RU");
|
||||||
}
|
}
|
||||||
|
|
||||||
onMounted(() => {
|
function kindLabel(kind: string | undefined): string {
|
||||||
loadProfiles();
|
return kind === "llm" ? "llm" : "heuristic";
|
||||||
loadTargetFields();
|
}
|
||||||
|
|
||||||
|
function canSave(): boolean {
|
||||||
|
if (!profileName.value.trim()) return false;
|
||||||
|
if (profileKind.value === "llm") return !!llmSchemaJson.value.trim();
|
||||||
|
return !!profileJson.value.trim();
|
||||||
|
}
|
||||||
|
|
||||||
|
onMounted(async () => {
|
||||||
|
await Promise.all([loadProfiles(), loadTargetFields(), loadLlmDefaults()]);
|
||||||
});
|
});
|
||||||
</script>
|
</script>
|
||||||
|
|
||||||
@@ -228,9 +469,10 @@ onMounted(() => {
|
|||||||
<div class="page">
|
<div class="page">
|
||||||
<h2 class="page-heading">Профили парсера</h2>
|
<h2 class="page-heading">Профили парсера</h2>
|
||||||
<p class="intro">
|
<p class="intro">
|
||||||
Generate строит статичные правила по образцу поста. Если Preview пустой по полю —
|
Два вида профилей под фиксированные поля Event:
|
||||||
напишите подсказку и снова Generate (агент правит текущий JSON). Можно править JSON
|
<strong>heuristic</strong> — статичные правила (Generate один раз, runtime без LLM);
|
||||||
вручную. Дальше профиль подключается к каналу на
|
<strong>llm</strong> — instruction + schema, DeepSeek на каждый пост в batch.
|
||||||
|
Связка с каналом — на
|
||||||
<router-link to="/parsers">Парсеры</router-link>.
|
<router-link to="/parsers">Парсеры</router-link>.
|
||||||
</p>
|
</p>
|
||||||
|
|
||||||
@@ -246,6 +488,13 @@ onMounted(() => {
|
|||||||
Название
|
Название
|
||||||
<input v-model="profileName" type="text" placeholder="Сводка: дата + НП + coords" />
|
<input v-model="profileName" type="text" placeholder="Сводка: дата + НП + coords" />
|
||||||
</label>
|
</label>
|
||||||
|
<label>
|
||||||
|
Тип
|
||||||
|
<select v-model="profileKind" :disabled="editingId != null">
|
||||||
|
<option value="heuristic">heuristic (правила)</option>
|
||||||
|
<option value="llm">llm (DeepSeek runtime)</option>
|
||||||
|
</select>
|
||||||
|
</label>
|
||||||
<label>
|
<label>
|
||||||
Статус
|
Статус
|
||||||
<select v-model="profileStatus">
|
<select v-model="profileStatus">
|
||||||
@@ -263,35 +512,64 @@ onMounted(() => {
|
|||||||
placeholder="15.03.2024 Населённый пункт Описание… 48.123456, 37.654321"
|
placeholder="15.03.2024 Населённый пункт Описание… 48.123456, 37.654321"
|
||||||
/>
|
/>
|
||||||
</label>
|
</label>
|
||||||
<label>
|
|
||||||
Подсказка агенту
|
<template v-if="profileKind === 'heuristic'">
|
||||||
<span class="muted"> (если Generate ошибся — опишите, что исправить, и нажмите Generate снова)</span>
|
<label>
|
||||||
<textarea
|
Подсказка агенту
|
||||||
v-model="generateHint"
|
<span class="muted"> (если Generate ошибся — опишите, что исправить, и нажмите Generate снова)</span>
|
||||||
class="hint-area"
|
<textarea
|
||||||
rows="3"
|
v-model="generateHint"
|
||||||
placeholder="например: description — абзацы между заголовком и координатами, без #хештегов"
|
class="hint-area"
|
||||||
/>
|
rows="3"
|
||||||
</label>
|
placeholder="например: description — абзацы между заголовком и координатами, без #хештегов"
|
||||||
|
/>
|
||||||
|
</label>
|
||||||
|
</template>
|
||||||
|
|
||||||
|
<template v-else>
|
||||||
|
<label>
|
||||||
|
Instruction
|
||||||
|
<textarea
|
||||||
|
v-model="llmInstruction"
|
||||||
|
class="hint-area"
|
||||||
|
rows="3"
|
||||||
|
placeholder="Инструкция для DeepSeek на каждый пост"
|
||||||
|
/>
|
||||||
|
</label>
|
||||||
|
<label class="required-item multi-flag">
|
||||||
|
<input v-model="llmMultiEvent" type="checkbox" />
|
||||||
|
Несколько событий в посте
|
||||||
|
<span class="muted"> (LLM вернёт массив; ingest с URL #e1, #e2…)</span>
|
||||||
|
</label>
|
||||||
|
</template>
|
||||||
|
|
||||||
<div class="form-actions">
|
<div class="form-actions">
|
||||||
<button
|
<button
|
||||||
type="button"
|
type="button"
|
||||||
class="btn btn-primary"
|
class="btn btn-primary"
|
||||||
:disabled="generating || !samplePost.trim()"
|
:disabled="generating || (profileKind === 'heuristic' && !samplePost.trim())"
|
||||||
@click="handleGenerate"
|
@click="handleGenerate"
|
||||||
>
|
>
|
||||||
{{
|
<template v-if="profileKind === 'llm'">
|
||||||
generating
|
{{ generating ? "…" : "Подставить дефолты schema" }}
|
||||||
? "Генерация..."
|
</template>
|
||||||
: generateHint.trim() || profileJson.trim()
|
<template v-else>
|
||||||
? "Generate / Refine"
|
{{
|
||||||
: "Generate"
|
generating
|
||||||
}}
|
? "Генерация..."
|
||||||
|
: generateHint.trim() || profileJson.trim()
|
||||||
|
? "Generate / Refine"
|
||||||
|
: "Generate"
|
||||||
|
}}
|
||||||
|
</template>
|
||||||
</button>
|
</button>
|
||||||
<button
|
<button
|
||||||
type="button"
|
type="button"
|
||||||
class="btn"
|
class="btn"
|
||||||
:disabled="previewing || !profileJson.trim()"
|
:disabled="
|
||||||
|
previewing ||
|
||||||
|
(profileKind === 'heuristic' ? !profileJson.trim() : !llmSchemaJson.trim())
|
||||||
|
"
|
||||||
@click="handlePreview"
|
@click="handlePreview"
|
||||||
>
|
>
|
||||||
{{ previewing ? "Preview..." : "Preview" }}
|
{{ previewing ? "Preview..." : "Preview" }}
|
||||||
@@ -299,7 +577,7 @@ onMounted(() => {
|
|||||||
<button
|
<button
|
||||||
type="button"
|
type="button"
|
||||||
class="btn btn-primary"
|
class="btn btn-primary"
|
||||||
:disabled="submitting || !profileJson.trim() || !profileName.trim()"
|
:disabled="submitting || !canSave()"
|
||||||
@click="handleSave"
|
@click="handleSave"
|
||||||
>
|
>
|
||||||
{{ submitting ? "Сохранение..." : editingId != null ? "Обновить" : "Сохранить профиль" }}
|
{{ submitting ? "Сохранение..." : editingId != null ? "Обновить" : "Сохранить профиль" }}
|
||||||
@@ -332,43 +610,120 @@ onMounted(() => {
|
|||||||
</section>
|
</section>
|
||||||
|
|
||||||
<section class="card">
|
<section class="card">
|
||||||
<h3>Профиль (JSON)</h3>
|
<h3 v-if="profileKind === 'heuristic'">Профиль (JSON)</h3>
|
||||||
|
<h3 v-else>extract_schema (JSON)</h3>
|
||||||
<textarea
|
<textarea
|
||||||
|
v-if="profileKind === 'heuristic'"
|
||||||
v-model="profileJson"
|
v-model="profileJson"
|
||||||
class="profile-area"
|
class="profile-area"
|
||||||
rows="16"
|
rows="16"
|
||||||
spellcheck="false"
|
spellcheck="false"
|
||||||
placeholder="Появится после Generate"
|
placeholder="Появится после Generate"
|
||||||
/>
|
/>
|
||||||
|
<textarea
|
||||||
|
v-else
|
||||||
|
v-model="llmSchemaJson"
|
||||||
|
class="profile-area"
|
||||||
|
rows="16"
|
||||||
|
spellcheck="false"
|
||||||
|
placeholder="Ключ → описание поля для LLM"
|
||||||
|
/>
|
||||||
</section>
|
</section>
|
||||||
</div>
|
</div>
|
||||||
|
|
||||||
<section v-if="previewFields" class="card">
|
<section v-if="previewFields" class="card">
|
||||||
<h3>Preview строки</h3>
|
<h3>
|
||||||
|
Preview
|
||||||
|
<span v-if="previewEvents.length > 1" class="muted">
|
||||||
|
({{ previewEvents.length }} событий)
|
||||||
|
</span>
|
||||||
|
<span v-else>строки</span>
|
||||||
|
</h3>
|
||||||
|
<p v-if="profileKind === 'llm' && !previewIsEvent" class="form-warn">
|
||||||
|
is_event=false — пост не будет ингеститься.
|
||||||
|
</p>
|
||||||
|
<p v-if="profileKind === 'llm' && previewEvents.length > 1" class="form-success match-ok">
|
||||||
|
Из поста извлечено {{ previewEvents.length }} событий;
|
||||||
|
пройдут required_fields: {{ previewMatchedCount }}.
|
||||||
|
Runtime URL: пост#e1 … #e{{ previewEvents.length }}.
|
||||||
|
</p>
|
||||||
<p v-if="emptyFields.length" class="form-warn">
|
<p v-if="emptyFields.length" class="form-warn">
|
||||||
Пустые поля: <code>{{ emptyFields.join(", ") }}</code> — уточните подсказку и Generate / Refine.
|
Пустые поля (первая строка): <code>{{ emptyFields.join(", ") }}</code>
|
||||||
|
<template v-if="profileKind === 'heuristic'">
|
||||||
|
— уточните подсказку и Generate / Refine.
|
||||||
|
</template>
|
||||||
|
</p>
|
||||||
|
<p v-if="!requiredFields.length" class="form-warn">
|
||||||
|
Нет обязательных полей —
|
||||||
|
<template v-if="profileKind === 'llm'">
|
||||||
|
фильтр только по is_event.
|
||||||
|
</template>
|
||||||
|
<template v-else>
|
||||||
|
любой пост канала будет сохранён. Отметьте поля, без которых пост нужно отбрасывать.
|
||||||
|
</template>
|
||||||
|
</p>
|
||||||
|
<p v-else-if="!previewMatched && previewEvents.length <= 1" class="form-warn">
|
||||||
|
Этот образец не прошёл бы фильтр. Не хватает:
|
||||||
|
<code>{{ previewMissing.join(", ") }}</code>
|
||||||
|
</p>
|
||||||
|
<p v-else-if="previewMatched" class="form-success match-ok">
|
||||||
|
<template v-if="previewEvents.length <= 1">
|
||||||
|
Образец проходит фильтр. Runtime сохранит только посты с заполненными
|
||||||
|
<code>{{ requiredFields.join(", ") }}</code>.
|
||||||
|
</template>
|
||||||
|
<template v-else>
|
||||||
|
Обязательные поля: <code>{{ requiredFields.join(", ") }}</code>
|
||||||
|
(проверка на каждое событие).
|
||||||
|
</template>
|
||||||
</p>
|
</p>
|
||||||
<div class="table-wrap">
|
<div class="table-wrap">
|
||||||
<table class="admin-table">
|
<table class="admin-table">
|
||||||
<thead>
|
<thead>
|
||||||
<tr>
|
<tr>
|
||||||
|
<th v-if="previewEvents.length > 1">#</th>
|
||||||
<th v-for="key in Object.keys(previewFields)" :key="key">{{ key }}</th>
|
<th v-for="key in Object.keys(previewFields)" :key="key">{{ key }}</th>
|
||||||
</tr>
|
</tr>
|
||||||
</thead>
|
</thead>
|
||||||
<tbody>
|
<tbody>
|
||||||
<tr>
|
<tr v-for="(row, idx) in (previewEvents.length ? previewEvents : [previewFields])" :key="idx">
|
||||||
|
<td v-if="previewEvents.length > 1">{{ idx + 1 }}</td>
|
||||||
<td
|
<td
|
||||||
v-for="(val, key) in previewFields"
|
v-for="key in Object.keys(previewFields)"
|
||||||
:key="key"
|
:key="key"
|
||||||
class="preview-cell"
|
class="preview-cell"
|
||||||
:class="{ empty: !val }"
|
:class="{ empty: !row?.[key] }"
|
||||||
>
|
>
|
||||||
{{ val || "—" }}
|
{{ row?.[key] || "—" }}
|
||||||
</td>
|
</td>
|
||||||
</tr>
|
</tr>
|
||||||
</tbody>
|
</tbody>
|
||||||
</table>
|
</table>
|
||||||
</div>
|
</div>
|
||||||
|
<div class="required-box">
|
||||||
|
<h4>Обязательно для сохранения</h4>
|
||||||
|
<p class="muted required-hint">
|
||||||
|
<template v-if="profileKind === 'llm'">
|
||||||
|
После LLM: событие отбрасывается, если обязательное поле пустое (поверх is_event).
|
||||||
|
</template>
|
||||||
|
<template v-else>
|
||||||
|
Пост отбрасывается, если хоть одно отмеченное поле не извлеклось (для coords — ещё и не парсится lat,lon).
|
||||||
|
</template>
|
||||||
|
</p>
|
||||||
|
<div class="required-grid">
|
||||||
|
<label
|
||||||
|
v-for="key in Object.keys(previewFields)"
|
||||||
|
:key="key"
|
||||||
|
class="required-item"
|
||||||
|
>
|
||||||
|
<input
|
||||||
|
type="checkbox"
|
||||||
|
:checked="isRequired(key)"
|
||||||
|
@change="toggleRequired(key)"
|
||||||
|
/>
|
||||||
|
<code>{{ key }}</code>
|
||||||
|
</label>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
</section>
|
</section>
|
||||||
|
|
||||||
<p v-if="error" class="form-error">{{ error }}</p>
|
<p v-if="error" class="form-error">{{ error }}</p>
|
||||||
@@ -382,6 +737,7 @@ onMounted(() => {
|
|||||||
<tr>
|
<tr>
|
||||||
<th>ID</th>
|
<th>ID</th>
|
||||||
<th>Название</th>
|
<th>Название</th>
|
||||||
|
<th>Тип</th>
|
||||||
<th>Статус</th>
|
<th>Статус</th>
|
||||||
<th>Создан</th>
|
<th>Создан</th>
|
||||||
<th></th>
|
<th></th>
|
||||||
@@ -391,6 +747,7 @@ onMounted(() => {
|
|||||||
<tr v-for="p in profiles" :key="p.id">
|
<tr v-for="p in profiles" :key="p.id">
|
||||||
<td>{{ p.id }}</td>
|
<td>{{ p.id }}</td>
|
||||||
<td>{{ p.name }}</td>
|
<td>{{ p.name }}</td>
|
||||||
|
<td><code>{{ kindLabel(p.kind) }}</code></td>
|
||||||
<td>{{ p.status }}</td>
|
<td>{{ p.status }}</td>
|
||||||
<td>{{ formatDate(p.created_at) }}</td>
|
<td>{{ formatDate(p.created_at) }}</td>
|
||||||
<td class="actions">
|
<td class="actions">
|
||||||
@@ -401,7 +758,7 @@ onMounted(() => {
|
|||||||
</td>
|
</td>
|
||||||
</tr>
|
</tr>
|
||||||
<tr v-if="!loading && !profiles.length">
|
<tr v-if="!loading && !profiles.length">
|
||||||
<td colspan="5" class="muted">Пока нет профилей</td>
|
<td colspan="6" class="muted">Пока нет профилей</td>
|
||||||
</tr>
|
</tr>
|
||||||
</tbody>
|
</tbody>
|
||||||
</table>
|
</table>
|
||||||
@@ -486,6 +843,42 @@ onMounted(() => {
|
|||||||
margin: 0 0 0.75rem;
|
margin: 0 0 0.75rem;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
.match-ok {
|
||||||
|
margin: 0 0 0.75rem;
|
||||||
|
}
|
||||||
|
|
||||||
|
.required-box {
|
||||||
|
margin-top: 1rem;
|
||||||
|
}
|
||||||
|
|
||||||
|
.required-box h4 {
|
||||||
|
margin: 0 0 0.25rem;
|
||||||
|
font-size: 0.9rem;
|
||||||
|
}
|
||||||
|
|
||||||
|
.required-hint {
|
||||||
|
margin: 0 0 0.6rem;
|
||||||
|
font-size: 0.8rem;
|
||||||
|
}
|
||||||
|
|
||||||
|
.required-grid {
|
||||||
|
display: flex;
|
||||||
|
flex-wrap: wrap;
|
||||||
|
gap: 0.5rem 1rem;
|
||||||
|
}
|
||||||
|
|
||||||
|
.required-item {
|
||||||
|
display: flex;
|
||||||
|
align-items: center;
|
||||||
|
gap: 0.35rem;
|
||||||
|
font-size: 0.85rem;
|
||||||
|
}
|
||||||
|
|
||||||
|
.multi-flag {
|
||||||
|
margin: 0.75rem 0 0;
|
||||||
|
flex-wrap: wrap;
|
||||||
|
}
|
||||||
|
|
||||||
.actions {
|
.actions {
|
||||||
display: flex;
|
display: flex;
|
||||||
gap: 0.35rem;
|
gap: 0.35rem;
|
||||||
|
|||||||
@@ -38,7 +38,11 @@ const editForm = ref({
|
|||||||
});
|
});
|
||||||
|
|
||||||
const readyProfiles = computed(() =>
|
const readyProfiles = computed(() =>
|
||||||
profiles.value.filter((p) => !!p.heuristic_profile),
|
profiles.value.filter((p) => {
|
||||||
|
if (p.status !== "ready") return false;
|
||||||
|
if (p.kind === "llm") return !!p.llm_profile;
|
||||||
|
return !!p.heuristic_profile;
|
||||||
|
}),
|
||||||
);
|
);
|
||||||
const activeChannels = computed(() => channels.value.filter((c) => c.is_active));
|
const activeChannels = computed(() => channels.value.filter((c) => c.is_active));
|
||||||
|
|
||||||
@@ -186,8 +190,17 @@ async function handleDelete(job: ParseJob) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
const STALE_JOB_MS = 15 * 60 * 1000;
|
||||||
|
|
||||||
|
function isStaleJob(job: ParseJob): boolean {
|
||||||
|
if (job.status !== "queued" && job.status !== "running") return false;
|
||||||
|
if (!job.last_run_at) return job.status === "running";
|
||||||
|
return Date.now() - new Date(job.last_run_at).getTime() >= STALE_JOB_MS;
|
||||||
|
}
|
||||||
|
|
||||||
function canForceRun(job: ParseJob): boolean {
|
function canForceRun(job: ParseJob): boolean {
|
||||||
return job.status !== "queued" && job.status !== "running";
|
if (job.status !== "queued" && job.status !== "running") return true;
|
||||||
|
return isStaleJob(job);
|
||||||
}
|
}
|
||||||
|
|
||||||
async function handleForceRun(jobId: number) {
|
async function handleForceRun(jobId: number) {
|
||||||
@@ -237,9 +250,10 @@ onUnmounted(() => {
|
|||||||
Создайте
|
Создайте
|
||||||
<router-link to="/channels">канал</router-link>
|
<router-link to="/channels">канал</router-link>
|
||||||
и
|
и
|
||||||
<router-link to="/parser-profiles">профиль</router-link>,
|
<router-link to="/parser-profiles">профиль</router-link>
|
||||||
затем свяжите их здесь. В Redis уходит плоский
|
(heuristic или llm), затем свяжите их здесь. Flatten в Redis:
|
||||||
<code>extract_mode=profile</code> + <code>heuristic_profile</code>.
|
heuristic → <code>extract_mode=profile</code>;
|
||||||
|
llm → <code>extract_mode=llm</code> (batch; listener для llm пока fallback).
|
||||||
</p>
|
</p>
|
||||||
<form class="admin-form" @submit.prevent="handleCreatePair">
|
<form class="admin-form" @submit.prevent="handleCreatePair">
|
||||||
<div class="form-row">
|
<div class="form-row">
|
||||||
@@ -257,7 +271,7 @@ onUnmounted(() => {
|
|||||||
<select v-model.number="pairForm.profile_id" required>
|
<select v-model.number="pairForm.profile_id" required>
|
||||||
<option :value="null" disabled>Выберите…</option>
|
<option :value="null" disabled>Выберите…</option>
|
||||||
<option v-for="p in readyProfiles" :key="p.id" :value="p.id">
|
<option v-for="p in readyProfiles" :key="p.id" :value="p.id">
|
||||||
{{ p.name }} (#{{ p.id }})
|
{{ p.name }} (#{{ p.id }}, {{ p.kind === "llm" ? "llm" : "heuristic" }})
|
||||||
</option>
|
</option>
|
||||||
</select>
|
</select>
|
||||||
</label>
|
</label>
|
||||||
@@ -343,7 +357,7 @@ onUnmounted(() => {
|
|||||||
<button
|
<button
|
||||||
class="btn btn-sm btn-danger"
|
class="btn btn-sm btn-danger"
|
||||||
type="button"
|
type="button"
|
||||||
:disabled="job.status === 'running'"
|
:disabled="job.status === 'running' && !isStaleJob(job)"
|
||||||
@click="handleDelete(job)"
|
@click="handleDelete(job)"
|
||||||
>
|
>
|
||||||
Удалить
|
Удалить
|
||||||
|
|||||||
@@ -54,14 +54,29 @@ async def run(job_id, source_config, *, ctx) -> tuple[list[dict], str | None]
|
|||||||
- `source_type` события = тип адаптера;
|
- `source_type` события = тип адаптера;
|
||||||
- `source_config` валидируется схемами из `contracts/sources.py`.
|
- `source_config` валидируется схемами из `contracts/sources.py`.
|
||||||
|
|
||||||
|
### Режимы извлечения (Telegram)
|
||||||
|
|
||||||
|
Цель всегда одна: фиксированные поля `IngestEventItem` / Event в ЦА (`title`, `description`, `locality`, `coords`→lat/lng, `event_date`, `topic`, …). Кастомные пользовательские таблицы — roadmap.
|
||||||
|
|
||||||
|
| `extract_mode` | Что делает | Откуда в UI |
|
||||||
|
|----------------|------------|-------------|
|
||||||
|
| `heuristic` | Legacy regex-парсер постов | raw job API (не пара) |
|
||||||
|
| `llm` | DeepSeek **на каждый** пост (batch) | профиль `kind=llm` → pair flatten |
|
||||||
|
| `profile` | Статичные правила `HeuristicProfile` | профиль `kind=heuristic` → pair flatten |
|
||||||
|
|
||||||
### LLM-режим (`extract_mode: llm`)
|
### LLM-режим (`extract_mode: llm`)
|
||||||
|
|
||||||
Для неструктурированных Telegram-постов и Crawl4AI:
|
Для неструктурированных Telegram-постов и Crawl4AI:
|
||||||
|
|
||||||
- ключ `DEEPSEEK_API_KEY` в `.env` (воркеры `cp-workers` / `cp-workers-web`);
|
- ключ `DEEPSEEK_API_KEY` в `.env` (воркеры `cp-workers` / `cp-workers-web`; preview в `ca-api`);
|
||||||
- Telegram: текст поста → DeepSeek JSON → `IngestEventItem`;
|
- Telegram batch: текст → DeepSeek JSON → один или несколько `IngestEventItem` (`workers/llm_extract.py`);
|
||||||
|
- опционально `extract_schema`, `instruction`, `required_fields` (post-extract gate поверх `is_event`);
|
||||||
|
- **`multi_event`** (opt-in в `llm_profile` / `TelegramSourceConfig`): модель возвращает массив `events`; каждое событие — отдельный ingest. URL: один event → `post.url`; несколько → `post.url#e1`, `#e2`, … (дедуп ЦА по `source_url`);
|
||||||
|
- heuristic / profile по-прежнему **1 пост → 1 событие**;
|
||||||
- Crawl4AI: страница → `LLMExtractionStrategy` (DeepSeek) с fallback на тот же DeepSeek по markdown;
|
- Crawl4AI: страница → `LLMExtractionStrategy` (DeepSeek) с fallback на тот же DeepSeek по markdown;
|
||||||
- в UI «Парсеры»: поле **Извлечение** = LLM DeepSeek.
|
- **listener:** для `llm` пока fallback на heuristic (LLM — batch-only by design).
|
||||||
|
|
||||||
|
Reusable LLM-профиль в ЦА: `ParserProfile.kind=llm` + JSON `llm_profile` (`contracts/llm_profile.py`). При enqueue flatten → `extract_mode=llm` + schema/instruction/required_fields/`multi_event`.
|
||||||
|
|
||||||
### Profile-режим (`extract_mode: profile`)
|
### Profile-режим (`extract_mode: profile`)
|
||||||
|
|
||||||
@@ -69,8 +84,12 @@ async def run(job_id, source_config, *, ctx) -> tuple[list[dict], str | None]
|
|||||||
|
|
||||||
- `heuristic_profile` в `source_config` (схема `contracts/heuristic_profile.py`);
|
- `heuristic_profile` в `source_config` (схема `contracts/heuristic_profile.py`);
|
||||||
- интерпретатор: `workers/heuristic_profile.py` (те же правила, что preview в ЦА);
|
- интерпретатор: `workers/heuristic_profile.py` (те же правила, что preview в ЦА);
|
||||||
- Telegram batch + listener применяют профиль без DeepSeek;
|
- `required_fields`: пост без заполненных обязательных полей не ингестится (batch + listener);
|
||||||
- генерация профиля — только в админке (`/admin/parser-profiles/generate`).
|
- генерация правил — один раз в админке (`/admin/parser-profiles/generate`); runtime **без** LLM.
|
||||||
|
|
||||||
|
Reusable heuristic-профиль: `ParserProfile.kind=heuristic` + `heuristic_profile`. Pair на «Парсеры» → flatten `extract_mode=profile`.
|
||||||
|
|
||||||
|
UI «Профили»: выбор `kind` (heuristic | llm). Связка канал+профиль на «Парсеры».
|
||||||
|
|
||||||
```mermaid
|
```mermaid
|
||||||
flowchart LR
|
flowchart LR
|
||||||
@@ -127,7 +146,10 @@ Legacy-ключ `cp:jobs` по-прежнему дренируется telegram-
|
|||||||
|
|
||||||
## Telegram real-time
|
## Telegram real-time
|
||||||
|
|
||||||
`TelegramListener` без изменений: подписки из ЦА, ingest с `listener: true` (статус `ParseJob` не трогается).
|
`TelegramListener`: подписки из ЦА, ingest с `listener: true` (статус `ParseJob` не трогается).
|
||||||
|
|
||||||
|
- `extract_mode=profile` — те же heuristic-правила, что batch;
|
||||||
|
- `extract_mode=llm` — **не** вызывает DeepSeek; fallback на legacy heuristic (LLM только в batch).
|
||||||
|
|
||||||
## Переменные окружения
|
## Переменные окружения
|
||||||
|
|
||||||
|
|||||||
@@ -11,9 +11,11 @@ from workers.heuristic_profile import extract_with_profile
|
|||||||
from workers.llm_extract import (
|
from workers.llm_extract import (
|
||||||
DEFAULT_EXTRACT_SCHEMA,
|
DEFAULT_EXTRACT_SCHEMA,
|
||||||
DEFAULT_INSTRUCTION,
|
DEFAULT_INSTRUCTION,
|
||||||
extract_event_fields,
|
event_source_url,
|
||||||
|
extract_event_list,
|
||||||
fields_to_ingest,
|
fields_to_ingest,
|
||||||
llm_enabled,
|
llm_enabled,
|
||||||
|
match_llm_required,
|
||||||
)
|
)
|
||||||
from workers.parsers.telegram_events import parse_event_posts
|
from workers.parsers.telegram_events import parse_event_posts
|
||||||
from workers.sources.telegram_client import (
|
from workers.sources.telegram_client import (
|
||||||
@@ -73,54 +75,72 @@ def _extract_posts_with_profile(posts, cfg: TelegramSourceConfig) -> tuple[list[
|
|||||||
text = (post.text or "").strip()
|
text = (post.text or "").strip()
|
||||||
if not text:
|
if not text:
|
||||||
continue
|
continue
|
||||||
events.append(
|
event = extract_with_profile(
|
||||||
extract_with_profile(
|
text,
|
||||||
text,
|
profile,
|
||||||
profile,
|
source_url=post.url,
|
||||||
source_url=post.url,
|
source_type="telegram",
|
||||||
source_type="telegram",
|
extra_metadata={
|
||||||
extra_metadata={
|
"channel": post.channel,
|
||||||
"channel": post.channel,
|
"message_id": post.id,
|
||||||
"message_id": post.id,
|
"post_date": post.date.isoformat() if post.date else None,
|
||||||
"post_date": post.date.isoformat() if post.date else None,
|
},
|
||||||
},
|
|
||||||
)
|
|
||||||
)
|
)
|
||||||
|
if event is None:
|
||||||
|
logger.debug("Profile skip %s (required fields missing)", post.url)
|
||||||
|
continue
|
||||||
|
events.append(event)
|
||||||
return events, None
|
return events, None
|
||||||
|
|
||||||
|
|
||||||
async def _extract_posts_with_llm(posts, cfg: TelegramSourceConfig) -> tuple[list[dict], str | None]:
|
async def _extract_posts_with_llm(posts, cfg: TelegramSourceConfig) -> tuple[list[dict], str | None]:
|
||||||
schema = cfg.extract_schema or DEFAULT_EXTRACT_SCHEMA
|
schema = cfg.extract_schema or DEFAULT_EXTRACT_SCHEMA
|
||||||
instruction = cfg.instruction or DEFAULT_INSTRUCTION
|
instruction = cfg.instruction or DEFAULT_INSTRUCTION
|
||||||
|
multi_event = bool(cfg.multi_event)
|
||||||
events: list[dict] = []
|
events: list[dict] = []
|
||||||
errors: list[str] = []
|
errors: list[str] = []
|
||||||
|
required = list(cfg.required_fields or [])
|
||||||
|
|
||||||
for post in posts:
|
for post in posts:
|
||||||
text = (post.text or "").strip()
|
text = (post.text or "").strip()
|
||||||
if not text:
|
if not text:
|
||||||
continue
|
continue
|
||||||
try:
|
try:
|
||||||
fields = await extract_event_fields(
|
field_list = await extract_event_list(
|
||||||
text,
|
text,
|
||||||
extract_schema=schema,
|
extract_schema=schema,
|
||||||
instruction=instruction,
|
instruction=instruction,
|
||||||
|
multi_event=multi_event,
|
||||||
)
|
)
|
||||||
if not fields.get("is_event", True):
|
if not field_list:
|
||||||
continue
|
continue
|
||||||
events.append(
|
total = len(field_list)
|
||||||
fields_to_ingest(
|
for idx, fields in enumerate(field_list):
|
||||||
source_type="telegram",
|
if not match_llm_required(fields, required):
|
||||||
source_url=post.url,
|
logger.debug(
|
||||||
raw_text=text,
|
"LLM skip %s event %s/%s (required fields missing)",
|
||||||
fields=fields,
|
post.url,
|
||||||
domain_profile="telegram_llm",
|
idx + 1,
|
||||||
extra_metadata={
|
total,
|
||||||
"channel": post.channel,
|
)
|
||||||
"message_id": post.id,
|
continue
|
||||||
"post_date": post.date.isoformat() if post.date else None,
|
events.append(
|
||||||
},
|
fields_to_ingest(
|
||||||
|
source_type="telegram",
|
||||||
|
source_url=event_source_url(post.url, idx, total),
|
||||||
|
raw_text=text,
|
||||||
|
fields=fields,
|
||||||
|
domain_profile="telegram_llm",
|
||||||
|
extra_metadata={
|
||||||
|
"channel": post.channel,
|
||||||
|
"message_id": post.id,
|
||||||
|
"post_date": post.date.isoformat() if post.date else None,
|
||||||
|
"event_index": idx + 1,
|
||||||
|
"event_count": total,
|
||||||
|
"multi_event": multi_event and total > 1,
|
||||||
|
},
|
||||||
|
)
|
||||||
)
|
)
|
||||||
)
|
|
||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
logger.exception("LLM extract failed for %s", post.url)
|
logger.exception("LLM extract failed for %s", post.url)
|
||||||
errors.append(f"{post.url}: {exc}")
|
errors.append(f"{post.url}: {exc}")
|
||||||
|
|||||||
@@ -4,7 +4,7 @@ from __future__ import annotations
|
|||||||
|
|
||||||
from typing import Any
|
from typing import Any
|
||||||
|
|
||||||
from contracts.heuristic_profile import HeuristicProfile, apply_profile
|
from contracts.heuristic_profile import HeuristicProfile, match_profile
|
||||||
from workers.llm_extract import parse_coords, parse_date
|
from workers.llm_extract import parse_coords, parse_date
|
||||||
|
|
||||||
|
|
||||||
@@ -62,8 +62,10 @@ def extract_with_profile(
|
|||||||
source_url: str,
|
source_url: str,
|
||||||
source_type: str = "telegram",
|
source_type: str = "telegram",
|
||||||
extra_metadata: dict | None = None,
|
extra_metadata: dict | None = None,
|
||||||
) -> dict:
|
) -> dict | None:
|
||||||
fields = apply_profile(text, profile)
|
matched, fields, _missing = match_profile(text, profile)
|
||||||
|
if not matched:
|
||||||
|
return None
|
||||||
return profile_fields_to_ingest(
|
return profile_fields_to_ingest(
|
||||||
source_type=source_type,
|
source_type=source_type,
|
||||||
source_url=source_url,
|
source_url=source_url,
|
||||||
|
|||||||
@@ -5,30 +5,31 @@ from __future__ import annotations
|
|||||||
import json
|
import json
|
||||||
import logging
|
import logging
|
||||||
import os
|
import os
|
||||||
import re
|
|
||||||
from datetime import datetime, timezone
|
from datetime import datetime, timezone
|
||||||
from typing import Any
|
from typing import Any
|
||||||
|
|
||||||
import httpx
|
import httpx
|
||||||
|
|
||||||
|
from contracts.heuristic_profile import parse_coords
|
||||||
|
from contracts.llm_profile import (
|
||||||
|
DEFAULT_EXTRACT_SCHEMA,
|
||||||
|
DEFAULT_INSTRUCTION,
|
||||||
|
match_llm_required,
|
||||||
|
)
|
||||||
|
|
||||||
logger = logging.getLogger("cp-worker.llm")
|
logger = logging.getLogger("cp-worker.llm")
|
||||||
|
|
||||||
COORDS_RE = re.compile(r"(-?\d{1,3}\.\d+)\s*,\s*(-?\d{1,3}\.\d+)")
|
# Re-export for adapters that import from this module
|
||||||
|
__all__ = [
|
||||||
DEFAULT_EXTRACT_SCHEMA: dict[str, str] = {
|
"DEFAULT_EXTRACT_SCHEMA",
|
||||||
"title": "string — short event title",
|
"DEFAULT_INSTRUCTION",
|
||||||
"locality": "string — place / settlement name",
|
"event_source_url",
|
||||||
"event_date": "string — date as DD.MM.YYYY or YYYY-MM-DD if known",
|
"extract_event_fields",
|
||||||
"description": "string — concise event summary",
|
"extract_event_list",
|
||||||
"coords": "string — latitude, longitude if present else empty",
|
"fields_to_ingest",
|
||||||
"topic": "string — short topic tag",
|
"llm_enabled",
|
||||||
}
|
"match_llm_required",
|
||||||
|
]
|
||||||
DEFAULT_INSTRUCTION = (
|
|
||||||
"Extract structured military/news event fields from the text. "
|
|
||||||
"If the text is not an event, return is_event=false. "
|
|
||||||
"Respond with a single JSON object only."
|
|
||||||
)
|
|
||||||
|
|
||||||
|
|
||||||
def llm_enabled() -> bool:
|
def llm_enabled() -> bool:
|
||||||
@@ -43,29 +44,57 @@ def llm_settings() -> dict[str, str]:
|
|||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
async def extract_event_fields(
|
def event_source_url(post_url: str, index: int, total: int) -> str:
|
||||||
text: str,
|
"""Stable per-event URL for CA dedup. Single event keeps bare post URL."""
|
||||||
|
base = (post_url or "").strip()
|
||||||
|
if total <= 1:
|
||||||
|
return base
|
||||||
|
# index is 0-based; fragment uses 1-based #eN
|
||||||
|
return f"{base}#e{index + 1}"
|
||||||
|
|
||||||
|
|
||||||
|
def _normalize_fields(raw: dict[str, Any], schema: dict[str, str]) -> dict[str, str]:
|
||||||
|
return {key: str(raw.get(key) or "").strip() for key in schema}
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_llm_content(
|
||||||
|
content: str,
|
||||||
|
schema: dict[str, str],
|
||||||
*,
|
*,
|
||||||
extract_schema: dict[str, str] | None = None,
|
multi_event: bool,
|
||||||
instruction: str | None = None,
|
) -> tuple[bool, list[dict[str, str]]]:
|
||||||
) -> dict[str, Any]:
|
"""Return (is_event, list of field dicts)."""
|
||||||
"""Ask DeepSeek to fill schema fields from free text. Returns dict (+ is_event)."""
|
parsed = json.loads(content)
|
||||||
|
if not isinstance(parsed, dict):
|
||||||
|
return False, []
|
||||||
|
|
||||||
|
is_event = bool(parsed.get("is_event", True))
|
||||||
|
if not is_event:
|
||||||
|
return False, []
|
||||||
|
|
||||||
|
if multi_event and isinstance(parsed.get("events"), list):
|
||||||
|
events: list[dict[str, str]] = []
|
||||||
|
for item in parsed["events"]:
|
||||||
|
if isinstance(item, dict):
|
||||||
|
events.append(_normalize_fields(item, schema))
|
||||||
|
return True, events
|
||||||
|
|
||||||
|
# Single-event fallback (legacy fields / top-level keys)
|
||||||
|
fields_raw = parsed.get("fields") if isinstance(parsed.get("fields"), dict) else parsed
|
||||||
|
if not isinstance(fields_raw, dict):
|
||||||
|
fields_raw = {}
|
||||||
|
# Drop non-schema keys that confuse normalize when falling back to top-level
|
||||||
|
cleaned = {k: fields_raw.get(k) for k in schema}
|
||||||
|
return True, [_normalize_fields(cleaned, schema)]
|
||||||
|
|
||||||
|
|
||||||
|
async def _call_deepseek(user_prompt: str) -> str:
|
||||||
settings = llm_settings()
|
settings = llm_settings()
|
||||||
if not settings["api_key"]:
|
if not settings["api_key"]:
|
||||||
raise RuntimeError(
|
raise RuntimeError(
|
||||||
"DEEPSEEK_API_KEY is not set. Add it to .env for LLM extract_mode."
|
"DEEPSEEK_API_KEY is not set. Add it to .env for LLM extract_mode."
|
||||||
)
|
)
|
||||||
|
|
||||||
schema = extract_schema or DEFAULT_EXTRACT_SCHEMA
|
|
||||||
instr = instruction or DEFAULT_INSTRUCTION
|
|
||||||
schema_lines = "\n".join(f"- {k}: {v}" for k, v in schema.items())
|
|
||||||
user_prompt = (
|
|
||||||
f"{instr}\n\n"
|
|
||||||
f"Fields to extract:\n{schema_lines}\n\n"
|
|
||||||
'Return JSON: {"is_event": true|false, "fields": {<field>: <string>}}\n\n'
|
|
||||||
f"Text:\n{text[:12000]}"
|
|
||||||
)
|
|
||||||
|
|
||||||
payload = {
|
payload = {
|
||||||
"model": settings["model"],
|
"model": settings["model"],
|
||||||
"messages": [
|
"messages": [
|
||||||
@@ -95,14 +124,72 @@ async def extract_event_fields(
|
|||||||
response.raise_for_status()
|
response.raise_for_status()
|
||||||
data = response.json()
|
data = response.json()
|
||||||
|
|
||||||
content = data["choices"][0]["message"]["content"]
|
return data["choices"][0]["message"]["content"]
|
||||||
parsed = json.loads(content)
|
|
||||||
fields = parsed.get("fields") if isinstance(parsed.get("fields"), dict) else parsed
|
|
||||||
if not isinstance(fields, dict):
|
async def extract_event_list(
|
||||||
fields = {}
|
text: str,
|
||||||
# Normalize to strings for known keys
|
*,
|
||||||
result = {key: str(fields.get(key) or "").strip() for key in schema}
|
extract_schema: dict[str, str] | None = None,
|
||||||
result["is_event"] = bool(parsed.get("is_event", True))
|
instruction: str | None = None,
|
||||||
|
multi_event: bool = False,
|
||||||
|
) -> list[dict[str, Any]]:
|
||||||
|
"""Extract zero or more event field dicts from free text.
|
||||||
|
|
||||||
|
Each dict has schema keys as strings. Empty list if not an event / no items.
|
||||||
|
"""
|
||||||
|
schema = extract_schema or DEFAULT_EXTRACT_SCHEMA
|
||||||
|
instr = instruction or DEFAULT_INSTRUCTION
|
||||||
|
schema_lines = "\n".join(f"- {k}: {v}" for k, v in schema.items())
|
||||||
|
|
||||||
|
if multi_event:
|
||||||
|
user_prompt = (
|
||||||
|
f"{instr}\n\n"
|
||||||
|
"If the text describes multiple distinct events (different places, "
|
||||||
|
"coords, or dates), return one object per event in \"events\".\n"
|
||||||
|
f"Fields per event:\n{schema_lines}\n\n"
|
||||||
|
'Return JSON: {"is_event": true|false, "events": [{<field>: <string>}, ...]}\n'
|
||||||
|
"If there is no event, return is_event=false and events=[].\n\n"
|
||||||
|
f"Text:\n{text[:12000]}"
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
user_prompt = (
|
||||||
|
f"{instr}\n\n"
|
||||||
|
f"Fields to extract:\n{schema_lines}\n\n"
|
||||||
|
'Return JSON: {"is_event": true|false, "fields": {<field>: <string>}}\n\n'
|
||||||
|
f"Text:\n{text[:12000]}"
|
||||||
|
)
|
||||||
|
|
||||||
|
content = await _call_deepseek(user_prompt)
|
||||||
|
is_event, events = _parse_llm_content(content, schema, multi_event=multi_event)
|
||||||
|
if not is_event:
|
||||||
|
return []
|
||||||
|
return events
|
||||||
|
|
||||||
|
|
||||||
|
async def extract_event_fields(
|
||||||
|
text: str,
|
||||||
|
*,
|
||||||
|
extract_schema: dict[str, str] | None = None,
|
||||||
|
instruction: str | None = None,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
"""Ask DeepSeek to fill schema fields from free text. Returns dict (+ is_event).
|
||||||
|
|
||||||
|
Single-event API kept for crawl4ai and legacy callers.
|
||||||
|
"""
|
||||||
|
events = await extract_event_list(
|
||||||
|
text,
|
||||||
|
extract_schema=extract_schema,
|
||||||
|
instruction=instruction,
|
||||||
|
multi_event=False,
|
||||||
|
)
|
||||||
|
if not events:
|
||||||
|
schema = extract_schema or DEFAULT_EXTRACT_SCHEMA
|
||||||
|
empty = {key: "" for key in schema}
|
||||||
|
empty["is_event"] = False
|
||||||
|
return empty
|
||||||
|
result = dict(events[0])
|
||||||
|
result["is_event"] = True
|
||||||
return result
|
return result
|
||||||
|
|
||||||
|
|
||||||
@@ -146,13 +233,6 @@ def fields_to_ingest(
|
|||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
def parse_coords(raw: str) -> tuple[float | None, float | None]:
|
|
||||||
match = COORDS_RE.search(raw or "")
|
|
||||||
if not match:
|
|
||||||
return None, None
|
|
||||||
return float(match.group(1)), float(match.group(2))
|
|
||||||
|
|
||||||
|
|
||||||
def parse_date(raw: str) -> datetime | None:
|
def parse_date(raw: str) -> datetime | None:
|
||||||
if not raw:
|
if not raw:
|
||||||
return None
|
return None
|
||||||
|
|||||||
@@ -81,7 +81,7 @@ class TelegramListener:
|
|||||||
", ".join(sorted(channels)) or "(none)",
|
", ".join(sorted(channels)) or "(none)",
|
||||||
)
|
)
|
||||||
|
|
||||||
def _build_event(self, channel: str, post) -> dict:
|
def _build_event(self, channel: str, post) -> dict | None:
|
||||||
sub = self._channels.get(normalize_channel(channel)) or {}
|
sub = self._channels.get(normalize_channel(channel)) or {}
|
||||||
cfg = sub.get("source_config") or {}
|
cfg = sub.get("source_config") or {}
|
||||||
extract_mode = cfg.get("extract_mode") or "heuristic"
|
extract_mode = cfg.get("extract_mode") or "heuristic"
|
||||||
@@ -107,6 +107,9 @@ class TelegramListener:
|
|||||||
|
|
||||||
async def _ingest_post(self, channel: str, post) -> None:
|
async def _ingest_post(self, channel: str, post) -> None:
|
||||||
event = self._build_event(channel, post)
|
event = self._build_event(channel, post)
|
||||||
|
if event is None:
|
||||||
|
logger.debug("Listener skip %s (profile required fields missing)", post.url)
|
||||||
|
return
|
||||||
sub = self._channels.get(normalize_channel(channel)) or {}
|
sub = self._channels.get(normalize_channel(channel)) or {}
|
||||||
job_id = sub.get("job_id")
|
job_id = sub.get("job_id")
|
||||||
|
|
||||||
|
|||||||
@@ -17,6 +17,8 @@ TARGET_FIELDS: tuple[str, ...] = (
|
|||||||
"region",
|
"region",
|
||||||
)
|
)
|
||||||
|
|
||||||
|
COORDS_RE = re.compile(r"(-?\d{1,3}\.\d+)\s*,\s*(-?\d{1,3}\.\d+)")
|
||||||
|
|
||||||
Strategy = Literal["regex", "line", "after_marker", "between", "full_text", "literal"]
|
Strategy = Literal["regex", "line", "after_marker", "between", "full_text", "literal"]
|
||||||
|
|
||||||
|
|
||||||
@@ -45,8 +47,24 @@ class FieldRule(BaseModel):
|
|||||||
class HeuristicProfile(BaseModel):
|
class HeuristicProfile(BaseModel):
|
||||||
version: Literal[1] = 1
|
version: Literal[1] = 1
|
||||||
fields: dict[str, FieldRule] = Field(default_factory=dict)
|
fields: dict[str, FieldRule] = Field(default_factory=dict)
|
||||||
|
required_fields: list[str] = Field(default_factory=list)
|
||||||
notes: str = ""
|
notes: str = ""
|
||||||
|
|
||||||
|
@field_validator("required_fields")
|
||||||
|
@classmethod
|
||||||
|
def known_required(cls, value: list[str]) -> list[str]:
|
||||||
|
seen: list[str] = []
|
||||||
|
unknown: list[str] = []
|
||||||
|
for name in value:
|
||||||
|
if name not in TARGET_FIELDS:
|
||||||
|
unknown.append(name)
|
||||||
|
continue
|
||||||
|
if name not in seen:
|
||||||
|
seen.append(name)
|
||||||
|
if unknown:
|
||||||
|
raise ValueError(f"Unknown required_fields: {sorted(set(unknown))}")
|
||||||
|
return seen
|
||||||
|
|
||||||
@model_validator(mode="after")
|
@model_validator(mode="after")
|
||||||
def known_fields_only(self) -> "HeuristicProfile":
|
def known_fields_only(self) -> "HeuristicProfile":
|
||||||
unknown = set(self.fields) - set(TARGET_FIELDS)
|
unknown = set(self.fields) - set(TARGET_FIELDS)
|
||||||
@@ -163,3 +181,38 @@ def apply_profile(text: str, profile: HeuristicProfile | dict[str, Any]) -> dict
|
|||||||
for name, rule in profile.fields.items():
|
for name, rule in profile.fields.items():
|
||||||
result[name] = _apply_rule(text, rule)
|
result[name] = _apply_rule(text, rule)
|
||||||
return result
|
return result
|
||||||
|
|
||||||
|
|
||||||
|
def parse_coords(raw: str) -> tuple[float | None, float | None]:
|
||||||
|
match = COORDS_RE.search(raw or "")
|
||||||
|
if not match:
|
||||||
|
return None, None
|
||||||
|
return float(match.group(1)), float(match.group(2))
|
||||||
|
|
||||||
|
|
||||||
|
def _field_is_filled(name: str, value: str) -> bool:
|
||||||
|
if not (value or "").strip():
|
||||||
|
return False
|
||||||
|
if name == "coords":
|
||||||
|
lat, lng = parse_coords(value)
|
||||||
|
return lat is not None and lng is not None
|
||||||
|
return True
|
||||||
|
|
||||||
|
|
||||||
|
def match_profile(
|
||||||
|
text: str,
|
||||||
|
profile: HeuristicProfile | dict[str, Any],
|
||||||
|
) -> tuple[bool, dict[str, str], list[str]]:
|
||||||
|
"""Apply profile and report whether required_fields are filled.
|
||||||
|
|
||||||
|
Empty required_fields means no gate (legacy profiles ingest every post).
|
||||||
|
"""
|
||||||
|
if isinstance(profile, dict):
|
||||||
|
profile = HeuristicProfile.model_validate(profile)
|
||||||
|
fields = apply_profile(text, profile)
|
||||||
|
missing = [
|
||||||
|
name
|
||||||
|
for name in profile.required_fields
|
||||||
|
if not _field_is_filled(name, fields.get(name, ""))
|
||||||
|
]
|
||||||
|
return (not missing, fields, missing)
|
||||||
|
|||||||
@@ -0,0 +1,86 @@
|
|||||||
|
"""Reusable LLM extract profile (CA storage + flattened into TelegramSourceConfig)."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from pydantic import BaseModel, Field, field_validator, model_validator
|
||||||
|
|
||||||
|
from contracts.heuristic_profile import TARGET_FIELDS
|
||||||
|
|
||||||
|
# LLM defaults omit region (same as workers/llm_extract.DEFAULT_EXTRACT_SCHEMA).
|
||||||
|
LLM_SCHEMA_FIELDS: tuple[str, ...] = (
|
||||||
|
"title",
|
||||||
|
"locality",
|
||||||
|
"event_date",
|
||||||
|
"description",
|
||||||
|
"coords",
|
||||||
|
"topic",
|
||||||
|
)
|
||||||
|
|
||||||
|
DEFAULT_EXTRACT_SCHEMA: dict[str, str] = {
|
||||||
|
"title": "string — short event title",
|
||||||
|
"locality": "string — place / settlement name",
|
||||||
|
"event_date": "string — date as DD.MM.YYYY or YYYY-MM-DD if known",
|
||||||
|
"description": "string — concise event summary",
|
||||||
|
"coords": "string — latitude, longitude if present else empty",
|
||||||
|
"topic": "string — short topic tag",
|
||||||
|
}
|
||||||
|
|
||||||
|
DEFAULT_INSTRUCTION = (
|
||||||
|
"Extract structured military/news event fields from the text. "
|
||||||
|
"If the text is not an event, return is_event=false. "
|
||||||
|
"Respond with a single JSON object only."
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
class LlmProfile(BaseModel):
|
||||||
|
instruction: str | None = None
|
||||||
|
extract_schema: dict[str, str] = Field(default_factory=lambda: dict(DEFAULT_EXTRACT_SCHEMA))
|
||||||
|
required_fields: list[str] = Field(default_factory=list)
|
||||||
|
# When true, LLM may return multiple events per post (array "events")
|
||||||
|
multi_event: bool = False
|
||||||
|
|
||||||
|
@field_validator("instruction", mode="before")
|
||||||
|
@classmethod
|
||||||
|
def empty_instruction_to_none(cls, value: Any) -> Any:
|
||||||
|
if value is None:
|
||||||
|
return None
|
||||||
|
if isinstance(value, str) and not value.strip():
|
||||||
|
return None
|
||||||
|
return value
|
||||||
|
|
||||||
|
@field_validator("required_fields")
|
||||||
|
@classmethod
|
||||||
|
def known_required(cls, value: list[str]) -> list[str]:
|
||||||
|
seen: list[str] = []
|
||||||
|
unknown: list[str] = []
|
||||||
|
for name in value:
|
||||||
|
if name not in TARGET_FIELDS:
|
||||||
|
unknown.append(name)
|
||||||
|
continue
|
||||||
|
if name not in seen:
|
||||||
|
seen.append(name)
|
||||||
|
if unknown:
|
||||||
|
raise ValueError(f"Unknown required_fields: {sorted(set(unknown))}")
|
||||||
|
return seen
|
||||||
|
|
||||||
|
@model_validator(mode="after")
|
||||||
|
def known_schema_keys(self) -> "LlmProfile":
|
||||||
|
if not self.extract_schema:
|
||||||
|
raise ValueError("extract_schema must not be empty")
|
||||||
|
unknown = set(self.extract_schema) - set(TARGET_FIELDS)
|
||||||
|
if unknown:
|
||||||
|
raise ValueError(f"Unknown extract_schema fields: {sorted(unknown)}")
|
||||||
|
return self
|
||||||
|
|
||||||
|
|
||||||
|
def match_llm_required(fields: dict[str, Any], required_fields: list[str]) -> bool:
|
||||||
|
"""True if all required_fields are non-empty strings (empty required = no gate)."""
|
||||||
|
if not required_fields:
|
||||||
|
return True
|
||||||
|
for name in required_fields:
|
||||||
|
val = fields.get(name)
|
||||||
|
if val is None or not str(val).strip():
|
||||||
|
return False
|
||||||
|
return True
|
||||||
+34
-8
@@ -15,6 +15,10 @@ class TelegramSourceConfig(BaseModel):
|
|||||||
extract_schema: dict[str, str] | None = None
|
extract_schema: dict[str, str] | None = None
|
||||||
instruction: str | None = None
|
instruction: str | None = None
|
||||||
heuristic_profile: dict | None = None
|
heuristic_profile: dict | None = None
|
||||||
|
# Post-extract gate for extract_mode=llm (empty = only is_event filter)
|
||||||
|
required_fields: list[str] = Field(default_factory=list)
|
||||||
|
# Split one post into N events when extract_mode=llm (opt-in)
|
||||||
|
multi_event: bool = False
|
||||||
sample_post: str | None = None # audit / re-generate; not required at runtime
|
sample_post: str | None = None # audit / re-generate; not required at runtime
|
||||||
|
|
||||||
@field_validator("channel")
|
@field_validator("channel")
|
||||||
@@ -22,15 +26,37 @@ class TelegramSourceConfig(BaseModel):
|
|||||||
def strip_channel(cls, value: str) -> str:
|
def strip_channel(cls, value: str) -> str:
|
||||||
return value.strip()
|
return value.strip()
|
||||||
|
|
||||||
@model_validator(mode="after")
|
@field_validator("required_fields")
|
||||||
def profile_requires_rules(self) -> "TelegramSourceConfig":
|
@classmethod
|
||||||
if self.extract_mode != "profile":
|
def known_required(cls, value: list[str]) -> list[str]:
|
||||||
return self
|
from contracts.heuristic_profile import TARGET_FIELDS
|
||||||
if not self.heuristic_profile:
|
|
||||||
raise ValueError("heuristic_profile required when extract_mode=profile")
|
|
||||||
from contracts.heuristic_profile import HeuristicProfile
|
|
||||||
|
|
||||||
HeuristicProfile.model_validate(self.heuristic_profile)
|
seen: list[str] = []
|
||||||
|
unknown: list[str] = []
|
||||||
|
for name in value:
|
||||||
|
if name not in TARGET_FIELDS:
|
||||||
|
unknown.append(name)
|
||||||
|
continue
|
||||||
|
if name not in seen:
|
||||||
|
seen.append(name)
|
||||||
|
if unknown:
|
||||||
|
raise ValueError(f"Unknown required_fields: {sorted(set(unknown))}")
|
||||||
|
return seen
|
||||||
|
|
||||||
|
@model_validator(mode="after")
|
||||||
|
def validate_extract_mode(self) -> "TelegramSourceConfig":
|
||||||
|
if self.extract_mode == "profile":
|
||||||
|
if not self.heuristic_profile:
|
||||||
|
raise ValueError("heuristic_profile required when extract_mode=profile")
|
||||||
|
from contracts.heuristic_profile import HeuristicProfile
|
||||||
|
|
||||||
|
HeuristicProfile.model_validate(self.heuristic_profile)
|
||||||
|
elif self.extract_mode == "llm" and self.extract_schema is not None:
|
||||||
|
from contracts.heuristic_profile import TARGET_FIELDS
|
||||||
|
|
||||||
|
unknown = set(self.extract_schema) - set(TARGET_FIELDS)
|
||||||
|
if unknown:
|
||||||
|
raise ValueError(f"Unknown extract_schema fields: {sorted(unknown)}")
|
||||||
return self
|
return self
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
@@ -184,7 +184,12 @@ async def run(job_id, source_config, *, ctx) -> tuple[list[dict], str | None]
|
|||||||
|
|
||||||
Опционально: `extract_mode: llm` (DeepSeek) для telegram/crawl4ai в batch.
|
Опционально: `extract_mode: llm` (DeepSeek) для telegram/crawl4ai в batch.
|
||||||
|
|
||||||
**Профиль + канал:** UI `/parser-profiles` (Generate/Preview/CRUD) и `/channels`; связка на `/parsers` → `ParseJob` с FK. При enqueue ЦА flatten в `extract_mode=profile` + `heuristic_profile`. Runtime (batch + listener) применяет только профиль, без LLM. Кастомные пользовательские таблицы — roadmap.
|
**Профиль + канал:** UI `/parser-profiles` (CRUD, Generate/Preview для heuristic; instruction/schema/Preview для llm) и `/channels`; связка на `/parsers` → `ParseJob` с FK. При enqueue ЦА flatten по `ParserProfile.kind`:
|
||||||
|
|
||||||
|
- `heuristic` → `extract_mode=profile` + `heuristic_profile` (runtime без LLM; listener поддерживает);
|
||||||
|
- `llm` → `extract_mode=llm` + `extract_schema` / `instruction` / `required_fields` (DeepSeek на каждый пост в batch; listener пока fallback на heuristic).
|
||||||
|
|
||||||
|
Целевые поля — фиксированная схема Event/`IngestEventItem`. Кастомные пользовательские таблицы — roadmap.
|
||||||
|
|
||||||
Подробности: [centers/parsing/ARCHITECTURE.md](../centers/parsing/ARCHITECTURE.md), [centers/analytics/ARCHITECTURE.md](../centers/analytics/ARCHITECTURE.md).
|
Подробности: [centers/parsing/ARCHITECTURE.md](../centers/parsing/ARCHITECTURE.md), [centers/analytics/ARCHITECTURE.md](../centers/analytics/ARCHITECTURE.md).
|
||||||
|
|
||||||
@@ -209,6 +214,7 @@ async def run(job_id, source_config, *, ctx) -> tuple[list[dict], str | None]
|
|||||||
| `contracts/sources.py` | Валидация `source_config` по `source_type` |
|
| `contracts/sources.py` | Валидация `source_config` по `source_type` |
|
||||||
| `contracts/queues.py` | `SOURCE_FAMILY` → ключ очереди |
|
| `contracts/queues.py` | `SOURCE_FAMILY` → ключ очереди |
|
||||||
| `contracts/heuristic_profile.py` | Статичный профиль конструктора + `apply_profile` |
|
| `contracts/heuristic_profile.py` | Статичный профиль конструктора + `apply_profile` |
|
||||||
|
| `contracts/llm_profile.py` | Reusable LLM instruction/schema + `match_llm_required` |
|
||||||
|
|
||||||
Правило: меняете форму события / конфиг источника / очередь — сначала `contracts/`, потом CA/CP/UI.
|
Правило: меняете форму события / конфиг источника / очередь — сначала `contracts/`, потом CA/CP/UI.
|
||||||
|
|
||||||
|
|||||||
+12
-3
@@ -13,7 +13,9 @@
|
|||||||
| [`jobs.py`](../contracts/jobs.py) | Контракт задания. Payload задания в Redis |
|
| [`jobs.py`](../contracts/jobs.py) | Контракт задания. Payload задания в Redis |
|
||||||
| [`ingest.py`](../contracts/ingest.py) |Контракт результата. Событие и пакет ingest |
|
| [`ingest.py`](../contracts/ingest.py) |Контракт результата. Событие и пакет ingest |
|
||||||
| [`sources.py`](../contracts/sources.py) | Контракт настроек парсера. Схемы `source_config` по `source_type` |
|
| [`sources.py`](../contracts/sources.py) | Контракт настроек парсера. Схемы `source_config` по `source_type` |
|
||||||
| [`queues.py`](../contracts/queues.py) | Контракт доставки задания нужному воркеру. `source_type` → family → Redis key |
|
| [`queues.py`](../contracts/queues.py) | Контракт доставки задания нужному воркеру. `source_type` → family → Redis key |
|
||||||
|
| [`heuristic_profile.py`](../contracts/heuristic_profile.py) | Статичные правила extract_mode=profile |
|
||||||
|
| [`llm_profile.py`](../contracts/llm_profile.py) | Reusable LLM instruction/schema + required_fields + `multi_event` |
|
||||||
|
|
||||||
ЦА при admin CRUD валидирует конфиг через `parse_source_config`.
|
ЦА при admin CRUD валидирует конфиг через `parse_source_config`.
|
||||||
ЦП адаптеры должны отдавать dict, совместимые с `IngestEventItem`.
|
ЦП адаптеры должны отдавать dict, совместимые с `IngestEventItem`.
|
||||||
@@ -57,13 +59,20 @@ Listener добавляет флаг `listener: true` на стороне CA API
|
|||||||
|
|
||||||
| source_type | Модель | Главные поля |
|
| source_type | Модель | Главные поля |
|
||||||
|-------------|--------|--------------|
|
|-------------|--------|--------------|
|
||||||
| `telegram` | `TelegramSourceConfig` | `channel`, `limit`, `extract_mode` (`heuristic`\|`llm`\|`profile`), `heuristic_profile` |
|
| `telegram` | `TelegramSourceConfig` | `channel`, `limit`, `extract_mode` (`heuristic`\|`llm`\|`profile`), `heuristic_profile`, `extract_schema`, `instruction`, `required_fields` (llm gate) |
|
||||||
| `crawl4ai` | `Crawl4AISourceConfig` | `urls`, `extract_mode`, `extract_schema`, `domain_profile` |
|
| `crawl4ai` | `Crawl4AISourceConfig` | `urls`, `extract_mode`, `extract_schema`, `domain_profile` |
|
||||||
| `viina` | `ViinaSourceConfig` | `urls` / `texts`, `input_mode` |
|
| `viina` | `ViinaSourceConfig` | `urls` / `texts`, `input_mode` |
|
||||||
|
|
||||||
Реестр: `CONFIG_MODELS` + `parse_source_config(source_type, raw)`.
|
Реестр: `CONFIG_MODELS` + `parse_source_config(source_type, raw)`.
|
||||||
|
|
||||||
Профиль парсера: `contracts/heuristic_profile.py` (`HeuristicProfile`, `apply_profile`). Сущности Profile/Channel живут в БД ЦА; в Redis уходит плоский `heuristic_profile`.
|
Профили в БД ЦА (`ParserProfile`):
|
||||||
|
|
||||||
|
| kind | Хранилище | Flatten в Redis |
|
||||||
|
|------|-----------|-----------------|
|
||||||
|
| `heuristic` | `heuristic_profile` (`contracts/heuristic_profile.py`) | `extract_mode=profile` |
|
||||||
|
| `llm` | `llm_profile` (`contracts/llm_profile.py`: instruction, extract_schema, required_fields, multi_event) | `extract_mode=llm` (+ `#eN` URLs if multi) |
|
||||||
|
|
||||||
|
`required_fields` — обязательные поля; пустой список = без фильтра (для llm остаётся только `is_event`). Канал + профиль живут в ЦА; CP получает только плоский `source_config`.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user