Add reusable LLM parser profiles with multi-event extract.

Support kind=llm profiles (instruction/schema), optional multi-event posts via #eN URLs, and recover stale running/queued parse jobs after worker crashes.

Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
2026-09-13 20:22:38 +03:00
co-authored by Cursor
parent 5811ecb134
commit 3f9dc6643b
27 changed files with 1334 additions and 205 deletions
+3
View File
@@ -27,6 +27,9 @@ ADMIN_JWT_SECRET=change-me-jwt-secret
# DEEPSEEK_BASE_URL=https://api.deepseek.com # DEEPSEEK_BASE_URL=https://api.deepseek.com
# DEEPSEEK_MODEL=deepseek-chat # DEEPSEEK_MODEL=deepseek-chat
# Recover parse jobs stuck in running/queued after worker crash (seconds, default 900)
# STALE_JOB_SECONDS=900
# CP adapter workers (set in docker-compose; override locally if needed) # CP adapter workers (set in docker-compose; override locally if needed)
# ENABLED_ADAPTERS=telegram # ENABLED_ADAPTERS=telegram
# WORKER_FAMILIES=telegram # WORKER_FAMILIES=telegram
+1 -1
View File
@@ -74,7 +74,7 @@ Adapters return dicts matching `IngestEventItem` (`contracts/ingest.py`). Requir
### LLM extract (runtime, not dev) ### LLM extract (runtime, not dev)
`extract_mode: llm` uses DeepSeek via `workers/llm_extract.py`. Key: `DEEPSEEK_API_KEY` in `.env`. Do not confuse with Cursor dev agents. `extract_mode: llm` uses DeepSeek via `workers/llm_extract.py` (batch). Reusable profiles in admin UI (`kind=heuristic|llm`) flatten into `extract_mode=profile` or `llm`. Key: `DEEPSEEK_API_KEY` in `.env`. Do not confuse with Cursor dev agents.
## Common tasks ## Common tasks
+7 -1
View File
@@ -76,7 +76,8 @@ docker compose up --build
| Раздел | Путь | Описание | | Раздел | Путь | Описание |
|--------|------|----------| |--------|------|----------|
| **Карта** | `/` | Интерактивная карта событий (публичный просмотр). CRUD ручных объектов — только для админа (ПКМ). Поддерживает `?eventId=` | | **Карта** | `/` | Интерактивная карта событий (публичный просмотр). CRUD ручных объектов — только для админа (ПКМ). Поддерживает `?eventId=` |
| **Парсеры** | `/parsers` | Адаптеры `telegram` / `crawl4ai` / `viina`, интервал, CRUD; дедуп по `source_url` | | **Парсеры** | `/parsers` | Связка канал + профиль (`heuristic`→`extract_mode=profile`, `llm`→`extract_mode=llm`); дедуп по `source_url` |
| **Профили** | `/parser-profiles` | Reusable heuristic (правила) или llm (instruction/schema); целевые поля = Event |
| **События** | `/events` | Фильтрация, пагинация, просмотр деталей, ссылка «На карте» для событий с координатами | | **События** | `/events` | Фильтрация, пагинация, просмотр деталей, ссылка «На карте» для событий с координатами |
| **Аналитика** | `/analytics` | KPI-карточки, график динамики ingest за 30 дней, топ населённых пунктов и регионов | | **Аналитика** | `/analytics` | KPI-карточки, график динамики ingest за 30 дней, топ населённых пунктов и регионов |
| **ПИ** | `/consumers` | CRUD подписчиков distribution API, ротация ключей, тест среза через `/api/v1/events` | | **ПИ** | `/consumers` | CRUD подписчиков distribution API, ротация ключей, тест среза через `/api/v1/events` |
@@ -214,3 +215,8 @@ docker compose down
``` ```
Данные PostgreSQL сохраняются в volume `pgdata`. Данные PostgreSQL сохраняются в volume `pgdata`.
## Вход в админку (/login):
Логин: admin
Пароль: change-me
+3
View File
@@ -98,8 +98,11 @@ class ParserProfile(Base):
id: Mapped[int] = mapped_column(Integer, primary_key=True, index=True) id: Mapped[int] = mapped_column(Integer, primary_key=True, index=True)
name: Mapped[str] = mapped_column(String(255), nullable=False) name: Mapped[str] = mapped_column(String(255), nullable=False)
# heuristic = static rules; llm = instruction + extract_schema at CP runtime
kind: Mapped[str] = mapped_column(String(50), default="heuristic", nullable=False, index=True)
sample_post: Mapped[str] = mapped_column(Text, default="") sample_post: Mapped[str] = mapped_column(Text, default="")
heuristic_profile: Mapped[dict | None] = mapped_column(JSON, nullable=True) heuristic_profile: Mapped[dict | None] = mapped_column(JSON, nullable=True)
llm_profile: Mapped[dict | None] = mapped_column(JSON, nullable=True)
status: Mapped[str] = mapped_column(String(50), default="draft", index=True) status: Mapped[str] = mapped_column(String(50), default="draft", index=True)
created_at: Mapped[datetime] = mapped_column( created_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), DateTime(timezone=True),
+16 -2
View File
@@ -87,7 +87,11 @@ def create_parse_job(payload: ParseJobCreate, db: Session = Depends(get_db)):
raise HTTPException(status_code=404, detail="Profile not found") raise HTTPException(status_code=404, detail="Profile not found")
if not channel: if not channel:
raise HTTPException(status_code=404, detail="Channel not found") raise HTTPException(status_code=404, detail="Channel not found")
if not profile.heuristic_profile: kind = (profile.kind or "heuristic").strip().lower()
if kind == "llm":
if not profile.llm_profile:
raise HTTPException(status_code=400, detail="LLM profile has no llm_profile")
elif not profile.heuristic_profile:
raise HTTPException(status_code=400, detail="Profile has no heuristic_profile") raise HTTPException(status_code=400, detail="Profile has no heuristic_profile")
if not channel.is_active: if not channel.is_active:
raise HTTPException(status_code=400, detail="Channel is inactive") raise HTTPException(status_code=400, detail="Channel is inactive")
@@ -146,7 +150,10 @@ def retry_parse_job(job_id: int, db: Session = Depends(get_db)):
job = db.query(ParseJob).filter(ParseJob.id == job_id).first() job = db.query(ParseJob).filter(ParseJob.id == job_id).first()
if not job: if not job:
raise HTTPException(status_code=404, detail="Job not found") raise HTTPException(status_code=404, detail="Job not found")
if job.status in ("queued", "running"):
from ..services.job_stale import is_stale_job
if job.status in ("queued", "running") and not is_stale_job(job):
raise HTTPException(status_code=409, detail="Job is already running or queued") raise HTTPException(status_code=409, detail="Job is already running or queued")
job.status = "queued" job.status = "queued"
@@ -218,7 +225,14 @@ def delete_parse_job(job_id: int, db: Session = Depends(get_db)):
if not job: if not job:
raise HTTPException(status_code=404, detail="Job not found") raise HTTPException(status_code=404, detail="Job not found")
if job.status == "running": if job.status == "running":
from ..services.job_stale import is_stale_job
if not is_stale_job(job):
raise HTTPException(status_code=409, detail="Cannot delete a running job") raise HTTPException(status_code=409, detail="Cannot delete a running job")
# Stale running — allow delete after marking failed for audit trail
job.status = "failed"
job.last_error = "Deleted while stale running"
db.commit()
db.delete(job) db.delete(job)
db.commit() db.commit()
@@ -52,7 +52,8 @@ def update_job_status(
raise HTTPException(status_code=404, detail="Job not found") raise HTTPException(status_code=404, detail="Job not found")
job.status = status job.status = status
if status in ("completed", "failed"): # Anchor staleness detection: running/queued start, and terminal finish
if status in ("running", "queued", "completed", "failed"):
job.last_run_at = datetime.now(timezone.utc) job.last_run_at = datetime.now(timezone.utc)
job.last_error = error job.last_error = error
db.commit() db.commit()
@@ -9,6 +9,7 @@ from pydantic import BaseModel, Field
from sqlalchemy.orm import Session from sqlalchemy.orm import Session
from contracts.heuristic_profile import HeuristicProfile from contracts.heuristic_profile import HeuristicProfile
from contracts.llm_profile import LlmProfile
from ..database import get_db from ..database import get_db
from ..deps import verify_admin from ..deps import verify_admin
@@ -44,6 +45,22 @@ class PreviewRequest(BaseModel):
class PreviewResponse(BaseModel): class PreviewResponse(BaseModel):
fields: dict[str, str] fields: dict[str, str]
matched: bool = True
missing_required: list[str] = Field(default_factory=list)
class PreviewLlmRequest(BaseModel):
sample_post: str = Field(min_length=1)
llm_profile: dict[str, Any]
class PreviewLlmResponse(BaseModel):
fields: dict[str, str]
events: list[dict[str, str]] = Field(default_factory=list)
matched: bool = True
missing_required: list[str] = Field(default_factory=list)
is_event: bool = True
matched_count: int = 0
def _validate_heuristic_profile(raw: dict[str, Any] | None) -> dict[str, Any] | None: def _validate_heuristic_profile(raw: dict[str, Any] | None) -> dict[str, Any] | None:
@@ -55,9 +72,33 @@ def _validate_heuristic_profile(raw: dict[str, Any] | None) -> dict[str, Any] |
raise HTTPException(status_code=400, detail=f"Invalid heuristic_profile: {exc}") from exc raise HTTPException(status_code=400, detail=f"Invalid heuristic_profile: {exc}") from exc
def _profile_status(heuristic_profile: dict | None, explicit: str | None = None) -> str: def _validate_llm_profile(raw: dict[str, Any] | None) -> dict[str, Any] | None:
if raw is None:
return None
try:
return LlmProfile.model_validate(raw).model_dump()
except Exception as exc:
raise HTTPException(status_code=400, detail=f"Invalid llm_profile: {exc}") from exc
def _normalize_kind(kind: str | None) -> str:
value = (kind or "heuristic").strip().lower()
if value not in ("heuristic", "llm"):
raise HTTPException(status_code=400, detail="kind must be heuristic or llm")
return value
def _profile_status(
*,
kind: str,
heuristic_profile: dict | None,
llm_profile: dict | None,
explicit: str | None = None,
) -> str:
if explicit: if explicit:
return explicit return explicit
if kind == "llm":
return "ready" if llm_profile else "draft"
return "ready" if heuristic_profile else "draft" return "ready" if heuristic_profile else "draft"
@@ -66,6 +107,18 @@ def target_fields():
return {"fields": builder.get_target_fields()} return {"fields": builder.get_target_fields()}
@router.get("/llm-defaults")
def llm_defaults():
from contracts.llm_profile import DEFAULT_EXTRACT_SCHEMA, DEFAULT_INSTRUCTION
return {
"instruction": DEFAULT_INSTRUCTION,
"extract_schema": DEFAULT_EXTRACT_SCHEMA,
"required_fields": [],
"multi_event": False,
}
@router.post("/generate", response_model=GenerateResponse) @router.post("/generate", response_model=GenerateResponse)
async def generate_profile(payload: GenerateRequest): async def generate_profile(payload: GenerateRequest):
if not builder.deepseek_enabled(): if not builder.deepseek_enabled():
@@ -87,6 +140,9 @@ async def generate_profile(payload: GenerateRequest):
detail=f"DeepSeek generate failed: {exc}", detail=f"DeepSeek generate failed: {exc}",
) from exc ) from exc
dumped = profile.model_dump() dumped = profile.model_dump()
if payload.current_profile and isinstance(payload.current_profile.get("required_fields"), list):
dumped["required_fields"] = payload.current_profile["required_fields"]
dumped = HeuristicProfile.model_validate(dumped).model_dump()
preview = builder.preview_with_profile(payload.sample_post, dumped) preview = builder.preview_with_profile(payload.sample_post, dumped)
return GenerateResponse( return GenerateResponse(
profile=dumped, profile=dumped,
@@ -102,10 +158,40 @@ def preview_profile(payload: PreviewRequest):
raise HTTPException(status_code=400, detail="heuristic_profile is required") raise HTTPException(status_code=400, detail="heuristic_profile is required")
try: try:
HeuristicProfile.model_validate(raw) HeuristicProfile.model_validate(raw)
fields = builder.preview_with_profile(payload.sample_post, raw) fields, matched, missing = builder.match_preview(payload.sample_post, raw)
except Exception as exc: except Exception as exc:
raise HTTPException(status_code=400, detail=str(exc)) from exc raise HTTPException(status_code=400, detail=str(exc)) from exc
return PreviewResponse(fields=fields) return PreviewResponse(fields=fields, matched=matched, missing_required=missing)
@router.post("/preview-llm", response_model=PreviewLlmResponse)
async def preview_llm_profile(payload: PreviewLlmRequest):
if not builder.deepseek_enabled():
raise HTTPException(
status_code=503,
detail="DEEPSEEK_API_KEY is not set on ca-api. Add it to .env for LLM preview.",
)
try:
LlmProfile.model_validate(payload.llm_profile)
events, fields, matched, missing, is_event, matched_count = await builder.preview_llm_extract(
payload.sample_post,
payload.llm_profile,
)
except ValueError as exc:
raise HTTPException(status_code=400, detail=str(exc)) from exc
except Exception as exc:
raise HTTPException(
status_code=502,
detail=f"DeepSeek LLM preview failed: {exc}",
) from exc
return PreviewLlmResponse(
fields=fields,
events=events,
matched=matched,
missing_required=missing,
is_event=is_event,
matched_count=matched_count,
)
@router.get("", response_model=list[ParserProfileRead]) @router.get("", response_model=list[ParserProfileRead])
@@ -115,12 +201,26 @@ def list_profiles(db: Session = Depends(get_db)):
@router.post("", response_model=ParserProfileRead, status_code=201) @router.post("", response_model=ParserProfileRead, status_code=201)
def create_profile(payload: ParserProfileCreate, db: Session = Depends(get_db)): def create_profile(payload: ParserProfileCreate, db: Session = Depends(get_db)):
kind = _normalize_kind(payload.kind)
heuristic = _validate_heuristic_profile(payload.heuristic_profile) heuristic = _validate_heuristic_profile(payload.heuristic_profile)
llm = _validate_llm_profile(payload.llm_profile)
if kind == "heuristic" and llm and not heuristic:
# ignore stray llm blob when creating heuristic
llm = None
if kind == "llm" and heuristic and not llm:
heuristic = None
profile = ParserProfile( profile = ParserProfile(
name=payload.name.strip(), name=payload.name.strip(),
kind=kind,
sample_post=payload.sample_post or "", sample_post=payload.sample_post or "",
heuristic_profile=heuristic, heuristic_profile=heuristic if kind == "heuristic" else None,
status=_profile_status(heuristic, payload.status), llm_profile=llm if kind == "llm" else None,
status=_profile_status(
kind=kind,
heuristic_profile=heuristic if kind == "heuristic" else None,
llm_profile=llm if kind == "llm" else None,
explicit=payload.status,
),
) )
db.add(profile) db.add(profile)
db.commit() db.commit()
@@ -150,12 +250,33 @@ def update_profile(
profile.name = payload.name.strip() profile.name = payload.name.strip()
if payload.sample_post is not None: if payload.sample_post is not None:
profile.sample_post = payload.sample_post profile.sample_post = payload.sample_post
if payload.kind is not None:
profile.kind = _normalize_kind(payload.kind)
kind = _normalize_kind(profile.kind)
if "heuristic_profile" in payload.model_fields_set: if "heuristic_profile" in payload.model_fields_set:
profile.heuristic_profile = _validate_heuristic_profile(payload.heuristic_profile) profile.heuristic_profile = _validate_heuristic_profile(payload.heuristic_profile)
if "llm_profile" in payload.model_fields_set:
profile.llm_profile = _validate_llm_profile(payload.llm_profile)
# Keep only the blob matching kind
if kind == "heuristic":
profile.llm_profile = None
else:
profile.heuristic_profile = None
if payload.status is not None: if payload.status is not None:
profile.status = payload.status profile.status = payload.status
elif "heuristic_profile" in payload.model_fields_set: elif (
profile.status = _profile_status(profile.heuristic_profile) "heuristic_profile" in payload.model_fields_set
or "llm_profile" in payload.model_fields_set
or payload.kind is not None
):
profile.status = _profile_status(
kind=kind,
heuristic_profile=profile.heuristic_profile,
llm_profile=profile.llm_profile,
)
db.commit() db.commit()
db.refresh(profile) db.refresh(profile)
+6
View File
@@ -161,15 +161,19 @@ class IngestResponse(BaseModel):
class ParserProfileCreate(BaseModel): class ParserProfileCreate(BaseModel):
name: str = Field(min_length=1, max_length=255) name: str = Field(min_length=1, max_length=255)
kind: str = Field(default="heuristic", pattern="^(heuristic|llm)$")
sample_post: str = "" sample_post: str = ""
heuristic_profile: dict[str, Any] | None = None heuristic_profile: dict[str, Any] | None = None
llm_profile: dict[str, Any] | None = None
status: str | None = None status: str | None = None
class ParserProfileUpdate(BaseModel): class ParserProfileUpdate(BaseModel):
name: str | None = Field(default=None, min_length=1, max_length=255) name: str | None = Field(default=None, min_length=1, max_length=255)
kind: str | None = Field(default=None, pattern="^(heuristic|llm)$")
sample_post: str | None = None sample_post: str | None = None
heuristic_profile: dict[str, Any] | None = None heuristic_profile: dict[str, Any] | None = None
llm_profile: dict[str, Any] | None = None
status: str | None = None status: str | None = None
@@ -178,8 +182,10 @@ class ParserProfileRead(BaseModel):
id: int id: int
name: str name: str
kind: str = "heuristic"
sample_post: str sample_post: str
heuristic_profile: dict[str, Any] | None heuristic_profile: dict[str, Any] | None
llm_profile: dict[str, Any] | None = None
status: str status: str
created_at: datetime created_at: datetime
@@ -0,0 +1,36 @@
"""Stale parse-job recovery helpers (running/queued left behind after worker crash)."""
from __future__ import annotations
import os
from datetime import datetime, timezone
from ..models import ParseJob
# Jobs stuck in running/queued longer than this are considered abandoned.
STALE_JOB_SECONDS = int(os.getenv("STALE_JOB_SECONDS", "900"))
def _aware(dt: datetime | None) -> datetime | None:
if dt is None:
return None
if dt.tzinfo is None:
return dt.replace(tzinfo=timezone.utc)
return dt
def job_anchor_time(job: ParseJob) -> datetime | None:
"""Best available timestamp for staleness (prefer last_run_at)."""
return _aware(job.last_run_at) or _aware(getattr(job, "created_at", None))
def is_stale_job(job: ParseJob, now: datetime | None = None, *, ttl: int | None = None) -> bool:
if job.status not in ("running", "queued"):
return False
now = now or datetime.now(timezone.utc)
anchor = job_anchor_time(job)
if anchor is None:
# No timestamp — treat long-lived running as stale immediately for recovery
return job.status == "running"
limit = ttl if ttl is not None else STALE_JOB_SECONDS
return (now - anchor).total_seconds() >= limit
@@ -33,6 +33,25 @@ def flatten_pair_config(
limit: int = 100, limit: int = 100,
) -> dict: ) -> dict:
"""Expand Profile + Channel into Redis/CP source_config.""" """Expand Profile + Channel into Redis/CP source_config."""
kind = (profile.kind or "heuristic").strip().lower()
if kind == "llm":
from contracts.llm_profile import LlmProfile
if not profile.llm_profile:
raise ValueError("ParserProfile.llm_profile is empty")
llm = LlmProfile.model_validate(profile.llm_profile)
cfg = TelegramSourceConfig(
channel=channel.channel.strip(),
limit=limit,
extract_mode="llm",
extract_schema=llm.extract_schema,
instruction=llm.instruction,
required_fields=list(llm.required_fields),
multi_event=bool(llm.multi_event),
sample_post=profile.sample_post or None,
)
return cfg.model_dump()
if not profile.heuristic_profile: if not profile.heuristic_profile:
raise ValueError("ParserProfile.heuristic_profile is empty") raise ValueError("ParserProfile.heuristic_profile is empty")
cfg = TelegramSourceConfig( cfg = TelegramSourceConfig(
@@ -23,11 +23,14 @@ def migrate_schema(engine: Engine) -> None:
if "channel_id" not in columns: if "channel_id" not in columns:
statements.append("ALTER TABLE parse_jobs ADD COLUMN channel_id INTEGER") statements.append("ALTER TABLE parse_jobs ADD COLUMN channel_id INTEGER")
# create_all handles new tables; FKs on existing DBs may need indexes if "parser_profiles" in tables:
if "parse_jobs" in tables: columns = {col["name"] for col in inspector.get_columns("parser_profiles")}
# Re-inspect after potential adds is not needed for FK constraints here — if "kind" not in columns:
# create_all + nullable FKs are enough for MVP; optional constraints below. statements.append(
pass "ALTER TABLE parser_profiles ADD COLUMN kind VARCHAR(50) NOT NULL DEFAULT 'heuristic'"
)
if "llm_profile" not in columns:
statements.append("ALTER TABLE parser_profiles ADD COLUMN llm_profile JSON")
if not statements: if not statements:
return return
@@ -14,6 +14,7 @@ from contracts.heuristic_profile import (
HeuristicProfile, HeuristicProfile,
TARGET_FIELDS, TARGET_FIELDS,
apply_profile, apply_profile,
match_profile,
target_field_specs, target_field_specs,
) )
@@ -47,6 +48,14 @@ def preview_with_profile(sample_post: str, profile: dict[str, Any] | HeuristicPr
return apply_profile(sample_post, profile) return apply_profile(sample_post, profile)
def match_preview(
sample_post: str,
profile: dict[str, Any] | HeuristicProfile,
) -> tuple[dict[str, str], bool, list[str]]:
matched, fields, missing = match_profile(sample_post, profile)
return fields, matched, missing
def empty_preview_fields(preview: dict[str, str]) -> list[str]: def empty_preview_fields(preview: dict[str, str]) -> list[str]:
return [name for name, value in preview.items() if not (value or "").strip()] return [name for name, value in preview.items() if not (value or "").strip()]
@@ -191,3 +200,111 @@ async def generate_profile(
content = data["choices"][0]["message"]["content"] content = data["choices"][0]["message"]["content"]
return _parse_profile_response(content) return _parse_profile_response(content)
async def preview_llm_extract(
sample_post: str,
llm_profile: dict[str, Any],
) -> tuple[list[dict[str, str]], dict[str, str], bool, list[str], bool, int]:
"""One-shot DeepSeek extract for CA preview (does not persist).
Returns (events, fields, matched, missing_required, is_event, matched_count).
``fields`` is the first event (or empty) for legacy UI compatibility.
"""
from contracts.llm_profile import DEFAULT_INSTRUCTION, LlmProfile, match_llm_required
settings = deepseek_settings()
if not settings["api_key"]:
raise RuntimeError(
"DEEPSEEK_API_KEY is not set on ca-api. Add it to .env for LLM preview."
)
sample = (sample_post or "").strip()
if not sample:
raise ValueError("sample_post is required")
profile = LlmProfile.model_validate(llm_profile)
schema = profile.extract_schema
instr = profile.instruction or DEFAULT_INSTRUCTION
schema_lines = "\n".join(f"- {k}: {v}" for k, v in schema.items())
multi = bool(profile.multi_event)
if multi:
user_prompt = (
f"{instr}\n\n"
"If the text describes multiple distinct events (different places, "
"coords, or dates), return one object per event in \"events\".\n"
f"Fields per event:\n{schema_lines}\n\n"
'Return JSON: {"is_event": true|false, "events": [{<field>: <string>}, ...]}\n'
"If there is no event, return is_event=false and events=[].\n\n"
f"Text:\n{sample[:12000]}"
)
else:
user_prompt = (
f"{instr}\n\n"
f"Fields to extract:\n{schema_lines}\n\n"
'Return JSON: {"is_event": true|false, "fields": {<field>: <string>}}\n\n'
f"Text:\n{sample[:12000]}"
)
payload = {
"model": settings["model"],
"messages": [
{
"role": "system",
"content": (
"You extract structured event data for a geoint map. "
"Output valid JSON only, no markdown."
),
},
{"role": "user", "content": user_prompt},
],
"temperature": 0.1,
"response_format": {"type": "json_object"},
}
url = f"{settings['base_url']}/chat/completions"
async with httpx.AsyncClient(timeout=90.0) as client:
response = await client.post(
url,
headers={
"Authorization": f"Bearer {settings['api_key']}",
"Content-Type": "application/json",
},
json=payload,
)
response.raise_for_status()
data = response.json()
content = data["choices"][0]["message"]["content"]
parsed = _extract_json_object(content)
is_event = bool(parsed.get("is_event", True))
empty_fields = {key: "" for key in schema}
if not is_event:
return [], empty_fields, False, [], False, 0
events: list[dict[str, str]] = []
if multi and isinstance(parsed.get("events"), list):
for item in parsed["events"]:
if isinstance(item, dict):
events.append({key: str(item.get(key) or "").strip() for key in schema})
else:
fields_raw = parsed.get("fields") if isinstance(parsed.get("fields"), dict) else parsed
if not isinstance(fields_raw, dict):
fields_raw = {}
events.append({key: str(fields_raw.get(key) or "").strip() for key in schema})
if not events:
return [], empty_fields, False, [], False, 0
matched_events = [ev for ev in events if match_llm_required(ev, list(profile.required_fields))]
fields = events[0]
missing = [
name
for name in profile.required_fields
if not str(fields.get(name) or "").strip()
]
matched = match_llm_required(fields, list(profile.required_fields))
return events, fields, matched, missing, True, len(matched_events)
@@ -5,6 +5,7 @@ from datetime import datetime, timezone
from ..database import SessionLocal from ..database import SessionLocal
from ..models import ParseJob from ..models import ParseJob
from .job_stale import STALE_JOB_SECONDS, is_stale_job
from .jobs import enqueue_parse_job from .jobs import enqueue_parse_job
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
@@ -13,10 +14,47 @@ TICK_SECONDS = 30
RECURRING_STATUSES = ("completed", "failed") RECURRING_STATUSES = ("completed", "failed")
def recover_stale_jobs(db, now: datetime) -> None:
"""Mark abandoned running/queued jobs as failed; re-queue if still active."""
stuck = (
db.query(ParseJob)
.filter(ParseJob.status.in_(("running", "queued")))
.all()
)
for job in stuck:
if not is_stale_job(job, now):
continue
prev = job.status
job.status = "failed"
job.last_error = (
f"Stale {prev} recovered after {STALE_JOB_SECONDS}s "
"(worker likely restarted)"
)
job.last_run_at = now
db.commit()
logger.warning("Recovered stale job %s (was %s)", job.id, prev)
if not job.is_active or job.interval_seconds <= 0:
continue
job.status = "queued"
job.last_error = None
db.commit()
try:
enqueue_parse_job(db, job)
logger.info("Re-queued recovered job %s", job.id)
except ValueError as exc:
job.status = "failed"
job.last_error = str(exc)
db.commit()
logger.warning("Skip re-queue recovered job %s: %s", job.id, exc)
def run_scheduler_tick() -> None: def run_scheduler_tick() -> None:
db = SessionLocal() db = SessionLocal()
try: try:
now = datetime.now(timezone.utc) now = datetime.now(timezone.utc)
recover_stale_jobs(db, now)
jobs = ( jobs = (
db.query(ParseJob) db.query(ParseJob)
.filter( .filter(
@@ -61,5 +99,9 @@ def start_scheduler() -> threading.Event:
stop_event = threading.Event() stop_event = threading.Event()
thread = threading.Thread(target=_scheduler_loop, args=(stop_event,), daemon=True) thread = threading.Thread(target=_scheduler_loop, args=(stop_event,), daemon=True)
thread.start() thread.start()
logger.info("Parse job scheduler started (tick every %ss)", TICK_SECONDS) logger.info(
"Parse job scheduler started (tick every %ss, stale after %ss)",
TICK_SECONDS,
STALE_JOB_SECONDS,
)
return stop_event return stop_event
+33 -3
View File
@@ -1,6 +1,7 @@
import { ADMIN_BASE, request } from "./client"; import { ADMIN_BASE, request } from "./client";
import type { import type {
HeuristicProfile, HeuristicProfile,
LlmProfile,
ParseChannel, ParseChannel,
ParseChannelCreate, ParseChannelCreate,
ParseChannelUpdate, ParseChannelUpdate,
@@ -10,7 +11,7 @@ import type {
TargetField, TargetField,
} from "../types/admin"; } from "../types/admin";
export type { HeuristicProfile, TargetField } from "../types/admin"; export type { HeuristicProfile, LlmProfile, TargetField } from "../types/admin";
export function fetchChannels(): Promise<ParseChannel[]> { export function fetchChannels(): Promise<ParseChannel[]> {
return request<ParseChannel[]>("/parse-channels", undefined, ADMIN_BASE); return request<ParseChannel[]>("/parse-channels", undefined, ADMIN_BASE);
@@ -64,6 +65,10 @@ export function fetchTargetFields(): Promise<{ fields: TargetField[] }> {
return request<{ fields: TargetField[] }>("/parser-profiles/target-fields", undefined, ADMIN_BASE); return request<{ fields: TargetField[] }>("/parser-profiles/target-fields", undefined, ADMIN_BASE);
} }
export function fetchLlmDefaults(): Promise<LlmProfile> {
return request<LlmProfile>("/parser-profiles/llm-defaults", undefined, ADMIN_BASE);
}
export function generateParserProfile(payload: { export function generateParserProfile(payload: {
sample_post: string; sample_post: string;
hint?: string; hint?: string;
@@ -94,10 +99,35 @@ export function generateParserProfile(payload: {
export function previewParserProfile( export function previewParserProfile(
sample_post: string, sample_post: string,
heuristic_profile: HeuristicProfile | Record<string, unknown>, heuristic_profile: HeuristicProfile | Record<string, unknown>,
): Promise<{ fields: Record<string, string> }> { ): Promise<{ fields: Record<string, string>; matched: boolean; missing_required: string[] }> {
return request<{ fields: Record<string, string> }>( return request<{ fields: Record<string, string>; matched: boolean; missing_required: string[] }>(
"/parser-profiles/preview", "/parser-profiles/preview",
{ method: "POST", body: JSON.stringify({ sample_post, heuristic_profile }) }, { method: "POST", body: JSON.stringify({ sample_post, heuristic_profile }) },
ADMIN_BASE, ADMIN_BASE,
); );
} }
export function previewLlmProfile(
sample_post: string,
llm_profile: LlmProfile | Record<string, unknown>,
): Promise<{
fields: Record<string, string>;
events: Record<string, string>[];
matched: boolean;
missing_required: string[];
is_event: boolean;
matched_count: number;
}> {
return request<{
fields: Record<string, string>;
events: Record<string, string>[];
matched: boolean;
missing_required: string[];
is_event: boolean;
matched_count: number;
}>(
"/parser-profiles/preview-llm",
{ method: "POST", body: JSON.stringify({ sample_post, llm_profile }) },
ADMIN_BASE,
);
}
+26 -12
View File
@@ -36,26 +36,52 @@ export interface ParseJobUpdate {
is_active?: boolean; is_active?: boolean;
} }
export type TargetField = {
name: string;
type: string;
description: string;
};
export type HeuristicProfile = {
version: number;
fields: Record<string, Record<string, unknown>>;
required_fields?: string[];
notes?: string;
};
export type LlmProfile = {
instruction?: string | null;
extract_schema: Record<string, string>;
required_fields?: string[];
multi_event?: boolean;
};
export interface ParserProfile { export interface ParserProfile {
id: number; id: number;
name: string; name: string;
kind: "heuristic" | "llm";
sample_post: string; sample_post: string;
heuristic_profile: Record<string, unknown> | null; heuristic_profile: Record<string, unknown> | null;
llm_profile: LlmProfile | Record<string, unknown> | null;
status: string; status: string;
created_at: string; created_at: string;
} }
export interface ParserProfileCreate { export interface ParserProfileCreate {
name: string; name: string;
kind?: "heuristic" | "llm";
sample_post?: string; sample_post?: string;
heuristic_profile?: Record<string, unknown> | null; heuristic_profile?: Record<string, unknown> | null;
llm_profile?: LlmProfile | Record<string, unknown> | null;
status?: string; status?: string;
} }
export interface ParserProfileUpdate { export interface ParserProfileUpdate {
name?: string; name?: string;
kind?: "heuristic" | "llm";
sample_post?: string; sample_post?: string;
heuristic_profile?: Record<string, unknown> | null; heuristic_profile?: Record<string, unknown> | null;
llm_profile?: LlmProfile | Record<string, unknown> | null;
status?: string; status?: string;
} }
@@ -82,18 +108,6 @@ export interface ParseChannelUpdate {
is_active?: boolean; is_active?: boolean;
} }
export type TargetField = {
name: string;
type: string;
description: string;
};
export type HeuristicProfile = {
version: number;
fields: Record<string, Record<string, unknown>>;
notes?: string;
};
export interface EventRecord { export interface EventRecord {
id: number; id: number;
source_type: string; source_type: string;
@@ -1,14 +1,17 @@
<script setup lang="ts"> <script setup lang="ts">
import { onMounted, ref } from "vue"; import { onMounted, ref, watch } from "vue";
import { import {
createProfile, createProfile,
deleteProfile, deleteProfile,
fetchLlmDefaults,
fetchProfiles, fetchProfiles,
fetchTargetFields, fetchTargetFields,
generateParserProfile, generateParserProfile,
previewLlmProfile,
previewParserProfile, previewParserProfile,
updateProfile, updateProfile,
type HeuristicProfile, type HeuristicProfile,
type LlmProfile,
type TargetField, type TargetField,
} from "../api/entities"; } from "../api/entities";
import type { ParserProfile } from "../types/admin"; import type { ParserProfile } from "../types/admin";
@@ -19,18 +22,29 @@ const error = ref("");
const success = ref(""); const success = ref("");
const submitting = ref(false); const submitting = ref(false);
const profileKind = ref<"heuristic" | "llm">("heuristic");
const samplePost = ref(""); const samplePost = ref("");
const generateHint = ref(""); const generateHint = ref("");
const profileName = ref(""); const profileName = ref("");
const targetFields = ref<TargetField[]>([]); const targetFields = ref<TargetField[]>([]);
const profileJson = ref(""); const profileJson = ref("");
const llmInstruction = ref("");
const llmSchemaJson = ref("");
const llmMultiEvent = ref(false);
const previewFields = ref<Record<string, string> | null>(null); const previewFields = ref<Record<string, string> | null>(null);
const previewEvents = ref<Record<string, string>[]>([]);
const emptyFields = ref<string[]>([]); const emptyFields = ref<string[]>([]);
const requiredFields = ref<string[]>([]);
const previewMatched = ref(true);
const previewMissing = ref<string[]>([]);
const previewIsEvent = ref(true);
const previewMatchedCount = ref(0);
const generating = ref(false); const generating = ref(false);
const previewing = ref(false); const previewing = ref(false);
const loadingFields = ref(true); const loadingFields = ref(true);
const editingId = ref<number | null>(null); const editingId = ref<number | null>(null);
const profileStatus = ref<"draft" | "ready">("draft"); const profileStatus = ref<"draft" | "ready">("draft");
const llmDefaults = ref<LlmProfile | null>(null);
async function loadProfiles() { async function loadProfiles() {
loading.value = true; loading.value = true;
@@ -56,7 +70,28 @@ async function loadTargetFields() {
} }
} }
function syncProfileFromJson(): HeuristicProfile | null { async function loadLlmDefaults() {
try {
llmDefaults.value = await fetchLlmDefaults();
} catch {
llmDefaults.value = null;
}
}
function applyLlmDefaults() {
const d = llmDefaults.value;
if (!d) return;
llmInstruction.value = d.instruction || "";
llmSchemaJson.value = JSON.stringify(d.extract_schema || {}, null, 2);
if (!requiredFields.value.length) {
requiredFields.value = Array.isArray(d.required_fields) ? [...d.required_fields] : [];
}
if (typeof d.multi_event === "boolean") {
llmMultiEvent.value = d.multi_event;
}
}
function syncHeuristicFromJson(): HeuristicProfile | null {
if (!profileJson.value.trim()) return null; if (!profileJson.value.trim()) return null;
try { try {
return JSON.parse(profileJson.value) as HeuristicProfile; return JSON.parse(profileJson.value) as HeuristicProfile;
@@ -66,7 +101,88 @@ function syncProfileFromJson(): HeuristicProfile | null {
} }
} }
function syncLlmFromForm(): LlmProfile | null {
let schema: Record<string, string>;
try {
schema = JSON.parse(llmSchemaJson.value || "{}") as Record<string, string>;
} catch {
error.value = "Некорректный JSON extract_schema";
return null;
}
if (!schema || typeof schema !== "object" || !Object.keys(schema).length) {
error.value = "extract_schema не должен быть пустым";
return null;
}
return {
instruction: llmInstruction.value.trim() || null,
extract_schema: schema,
required_fields: [...requiredFields.value],
multi_event: llmMultiEvent.value,
};
}
const DEFAULT_REQUIRED = ["event_date", "coords"];
function uniqueFields(names: string[]): string[] {
return names.filter((name, idx) => names.indexOf(name) === idx);
}
function writeRequiredToHeuristicJson(names: string[]) {
const current = syncHeuristicFromJson();
if (!current) return;
current.required_fields = names;
profileJson.value = JSON.stringify(current, null, 2);
}
function requiredFromHeuristic(
profile: HeuristicProfile,
preview: Record<string, string> | null,
): string[] {
if (Array.isArray(profile.required_fields)) {
return uniqueFields(profile.required_fields);
}
if (!preview) return [];
return DEFAULT_REQUIRED.filter((name) => !!(preview[name] || "").trim());
}
function isRequired(name: string): boolean {
return requiredFields.value.includes(name);
}
async function toggleRequired(name: string) {
const next = isRequired(name)
? requiredFields.value.filter((n) => n !== name)
: [...requiredFields.value, name];
requiredFields.value = next;
if (profileKind.value === "heuristic") {
writeRequiredToHeuristicJson(next);
}
if (previewFields.value && samplePost.value.trim()) {
await handlePreview();
}
}
watch(profileKind, (kind, prev) => {
if (kind === prev) return;
previewFields.value = null;
previewEvents.value = [];
emptyFields.value = [];
previewMatched.value = true;
previewMissing.value = [];
previewIsEvent.value = true;
previewMatchedCount.value = 0;
success.value = "";
if (kind === "llm" && !llmSchemaJson.value.trim()) {
applyLlmDefaults();
}
});
async function handleGenerate() { async function handleGenerate() {
if (profileKind.value === "llm") {
applyLlmDefaults();
success.value = "Подставлены дефолты LLM schema/instruction. Отредактируйте и Preview.";
return;
}
if (!samplePost.value.trim()) { if (!samplePost.value.trim()) {
error.value = "Вставьте образец поста"; error.value = "Вставьте образец поста";
return; return;
@@ -79,7 +195,7 @@ async function handleGenerate() {
try { try {
let current: HeuristicProfile | null = null; let current: HeuristicProfile | null = null;
if (profileJson.value.trim()) { if (profileJson.value.trim()) {
current = syncProfileFromJson(); current = syncHeuristicFromJson();
if (!current) return; if (!current) return;
} }
const data = await generateParserProfile({ const data = await generateParserProfile({
@@ -87,9 +203,18 @@ async function handleGenerate() {
hint: generateHint.value.trim() || undefined, hint: generateHint.value.trim() || undefined,
current_profile: current, current_profile: current,
}); });
const required =
current && Array.isArray(current.required_fields)
? uniqueFields(current.required_fields)
: DEFAULT_REQUIRED.filter((name) => !!(data.preview[name] || "").trim());
data.profile.required_fields = required;
profileJson.value = JSON.stringify(data.profile, null, 2); profileJson.value = JSON.stringify(data.profile, null, 2);
requiredFields.value = required;
previewFields.value = data.preview; previewFields.value = data.preview;
emptyFields.value = data.empty_fields || []; emptyFields.value = data.empty_fields || [];
const previewData = await previewParserProfile(samplePost.value, data.profile);
previewMatched.value = previewData.matched;
previewMissing.value = previewData.missing_required || [];
if (emptyFields.value.length) { if (emptyFields.value.length) {
success.value = success.value =
`Профиль обновлён, но пустые поля: ${emptyFields.value.join(", ")}. ` + `Профиль обновлён, но пустые поля: ${emptyFields.value.join(", ")}. ` +
@@ -109,11 +234,6 @@ async function handleGenerate() {
} }
async function handlePreview() { async function handlePreview() {
const current = syncProfileFromJson();
if (!current) {
if (!error.value) error.value = "Сначала сгенерируйте или вставьте профиль";
return;
}
if (!samplePost.value.trim()) { if (!samplePost.value.trim()) {
error.value = "Нужен образец поста для preview"; error.value = "Нужен образец поста для preview";
return; return;
@@ -121,11 +241,53 @@ async function handlePreview() {
previewing.value = true; previewing.value = true;
error.value = ""; error.value = "";
try { try {
if (profileKind.value === "llm") {
const llm = syncLlmFromForm();
if (!llm) return;
const data = await previewLlmProfile(samplePost.value, llm);
const rows =
Array.isArray(data.events) && data.events.length
? data.events
: data.fields
? [data.fields]
: [];
previewEvents.value = rows;
previewFields.value = rows[0] || data.fields || null;
emptyFields.value = previewFields.value
? Object.entries(previewFields.value)
.filter(([, v]) => !v)
.map(([k]) => k)
: [];
previewMatched.value = data.matched;
previewMissing.value = data.missing_required || [];
previewIsEvent.value = data.is_event;
previewMatchedCount.value = data.matched_count ?? rows.length;
if (!data.is_event) {
success.value = "Модель пометила текст как не-событие (is_event=false) — runtime пропустит.";
} else if (rows.length > 1) {
success.value = `Preview: ${rows.length} событий из поста (пройдут фильтр: ${previewMatchedCount.value}).`;
}
} else {
const current = syncHeuristicFromJson();
if (!current) {
if (!error.value) error.value = "Сначала сгенерируйте или вставьте профиль";
return;
}
const data = await previewParserProfile(samplePost.value, current); const data = await previewParserProfile(samplePost.value, current);
previewFields.value = data.fields; previewFields.value = data.fields;
emptyFields.value = Object.entries(data.fields) emptyFields.value = Object.entries(data.fields)
.filter(([, v]) => !v) .filter(([, v]) => !v)
.map(([k]) => k); .map(([k]) => k);
requiredFields.value = requiredFromHeuristic(current, data.fields);
if (!Array.isArray(current.required_fields)) {
writeRequiredToHeuristicJson(requiredFields.value);
}
previewMatched.value = data.matched;
previewMissing.value = data.missing_required || [];
previewIsEvent.value = true;
previewEvents.value = data.fields ? [data.fields] : [];
previewMatchedCount.value = data.matched ? 1 : 0;
}
} catch (err) { } catch (err) {
error.value = err instanceof Error ? err.message : "Ошибка preview"; error.value = err instanceof Error ? err.message : "Ошибка preview";
} finally { } finally {
@@ -135,29 +297,72 @@ async function handlePreview() {
function resetEditor() { function resetEditor() {
editingId.value = null; editingId.value = null;
profileKind.value = "heuristic";
profileName.value = ""; profileName.value = "";
samplePost.value = ""; samplePost.value = "";
generateHint.value = ""; generateHint.value = "";
profileJson.value = ""; profileJson.value = "";
llmInstruction.value = "";
llmSchemaJson.value = "";
llmMultiEvent.value = false;
profileStatus.value = "draft"; profileStatus.value = "draft";
previewFields.value = null; previewFields.value = null;
previewEvents.value = [];
emptyFields.value = []; emptyFields.value = [];
requiredFields.value = [];
previewMatched.value = true;
previewMissing.value = [];
previewIsEvent.value = true;
previewMatchedCount.value = 0;
success.value = ""; success.value = "";
} }
function openEdit(p: ParserProfile) { function openEdit(p: ParserProfile) {
editingId.value = p.id; editingId.value = p.id;
profileKind.value = p.kind === "llm" ? "llm" : "heuristic";
profileName.value = p.name; profileName.value = p.name;
samplePost.value = p.sample_post || ""; samplePost.value = p.sample_post || "";
generateHint.value = ""; generateHint.value = "";
profileStatus.value = p.status === "ready" ? "ready" : "draft";
previewFields.value = null;
previewEvents.value = [];
emptyFields.value = [];
previewMatched.value = true;
previewMissing.value = [];
previewIsEvent.value = true;
previewMatchedCount.value = 0;
success.value = "";
error.value = "";
if (p.kind === "llm") {
profileJson.value = "";
const lp = (p.llm_profile || {}) as LlmProfile;
llmInstruction.value = lp.instruction || "";
llmSchemaJson.value = JSON.stringify(lp.extract_schema || {}, null, 2);
llmMultiEvent.value = !!lp.multi_event;
requiredFields.value = Array.isArray(lp.required_fields)
? uniqueFields(lp.required_fields)
: [];
if (!llmSchemaJson.value.trim() || llmSchemaJson.value === "{}") {
applyLlmDefaults();
}
} else {
llmInstruction.value = "";
llmSchemaJson.value = "";
llmMultiEvent.value = false;
profileJson.value = p.heuristic_profile profileJson.value = p.heuristic_profile
? JSON.stringify(p.heuristic_profile, null, 2) ? JSON.stringify(p.heuristic_profile, null, 2)
: ""; : "";
profileStatus.value = p.status === "ready" ? "ready" : "draft"; const hp = p.heuristic_profile;
previewFields.value = null; if (hp && Array.isArray(hp.required_fields)) {
emptyFields.value = []; requiredFields.value = uniqueFields(hp.required_fields as string[]);
success.value = ""; } else if (hp) {
error.value = ""; requiredFields.value = [...DEFAULT_REQUIRED];
writeRequiredToHeuristicJson(requiredFields.value);
} else {
requiredFields.value = [];
}
}
window.scrollTo({ top: 0, behavior: "smooth" }); window.scrollTo({ top: 0, behavior: "smooth" });
} }
@@ -166,34 +371,61 @@ async function handleSave() {
error.value = "Укажите название профиля"; error.value = "Укажите название профиля";
return; return;
} }
const current = syncProfileFromJson();
if (!current) {
if (!error.value) error.value = "Нужен JSON профиля (Generate или вручную)";
return;
}
submitting.value = true; submitting.value = true;
error.value = ""; error.value = "";
success.value = ""; success.value = "";
try { try {
if (editingId.value != null) { if (profileKind.value === "llm") {
await updateProfile(editingId.value, { const llm = syncLlmFromForm();
if (!llm) return;
const payload = {
name: profileName.value.trim(), name: profileName.value.trim(),
kind: "llm" as const,
sample_post: samplePost.value,
llm_profile: llm,
heuristic_profile: null,
status: profileStatus.value,
};
if (editingId.value != null) {
await updateProfile(editingId.value, payload);
success.value = `LLM-профиль #${editingId.value} обновлён`;
} else {
const created = await createProfile({
...payload,
status: profileStatus.value === "draft" ? "draft" : "ready",
});
success.value = `LLM-профиль #${created.id} сохранён`;
editingId.value = created.id;
profileStatus.value = created.status === "ready" ? "ready" : "draft";
}
} else {
writeRequiredToHeuristicJson(requiredFields.value);
const current = syncHeuristicFromJson();
if (!current) {
if (!error.value) error.value = "Нужен JSON профиля (Generate или вручную)";
return;
}
const payload = {
name: profileName.value.trim(),
kind: "heuristic" as const,
sample_post: samplePost.value, sample_post: samplePost.value,
heuristic_profile: current, heuristic_profile: current,
llm_profile: null,
status: profileStatus.value, status: profileStatus.value,
}); };
if (editingId.value != null) {
await updateProfile(editingId.value, payload);
success.value = `Профиль #${editingId.value} обновлён`; success.value = `Профиль #${editingId.value} обновлён`;
} else { } else {
const created = await createProfile({ const created = await createProfile({
name: profileName.value.trim(), ...payload,
sample_post: samplePost.value,
heuristic_profile: current,
status: profileStatus.value === "draft" ? "draft" : "ready", status: profileStatus.value === "draft" ? "draft" : "ready",
}); });
success.value = `Профиль #${created.id} сохранён`; success.value = `Профиль #${created.id} сохранён`;
editingId.value = created.id; editingId.value = created.id;
profileStatus.value = created.status === "ready" ? "ready" : "draft"; profileStatus.value = created.status === "ready" ? "ready" : "draft";
} }
}
await loadProfiles(); await loadProfiles();
} catch (err) { } catch (err) {
error.value = err instanceof Error ? err.message : "Не удалось сохранить профиль"; error.value = err instanceof Error ? err.message : "Не удалось сохранить профиль";
@@ -218,9 +450,18 @@ function formatDate(value: string): string {
return new Date(value).toLocaleString("ru-RU"); return new Date(value).toLocaleString("ru-RU");
} }
onMounted(() => { function kindLabel(kind: string | undefined): string {
loadProfiles(); return kind === "llm" ? "llm" : "heuristic";
loadTargetFields(); }
function canSave(): boolean {
if (!profileName.value.trim()) return false;
if (profileKind.value === "llm") return !!llmSchemaJson.value.trim();
return !!profileJson.value.trim();
}
onMounted(async () => {
await Promise.all([loadProfiles(), loadTargetFields(), loadLlmDefaults()]);
}); });
</script> </script>
@@ -228,9 +469,10 @@ onMounted(() => {
<div class="page"> <div class="page">
<h2 class="page-heading">Профили парсера</h2> <h2 class="page-heading">Профили парсера</h2>
<p class="intro"> <p class="intro">
Generate строит статичные правила по образцу поста. Если Preview пустой по полю — Два вида профилей под фиксированные поля Event:
напишите подсказку и снова Generate (агент правит текущий JSON). Можно править JSON <strong>heuristic</strong> — статичные правила (Generate один раз, runtime без LLM);
вручную. Дальше профиль подключается к каналу на <strong>llm</strong> — instruction + schema, DeepSeek на каждый пост в batch.
Связка с каналом — на
<router-link to="/parsers">Парсеры</router-link>. <router-link to="/parsers">Парсеры</router-link>.
</p> </p>
@@ -246,6 +488,13 @@ onMounted(() => {
Название Название
<input v-model="profileName" type="text" placeholder="Сводка: дата + НП + coords" /> <input v-model="profileName" type="text" placeholder="Сводка: дата + НП + coords" />
</label> </label>
<label>
Тип
<select v-model="profileKind" :disabled="editingId != null">
<option value="heuristic">heuristic (правила)</option>
<option value="llm">llm (DeepSeek runtime)</option>
</select>
</label>
<label> <label>
Статус Статус
<select v-model="profileStatus"> <select v-model="profileStatus">
@@ -263,6 +512,8 @@ onMounted(() => {
placeholder="15.03.2024 Населённый пункт&#10;&#10;Описание…&#10;&#10;48.123456, 37.654321" placeholder="15.03.2024 Населённый пункт&#10;&#10;Описание…&#10;&#10;48.123456, 37.654321"
/> />
</label> </label>
<template v-if="profileKind === 'heuristic'">
<label> <label>
Подсказка агенту Подсказка агенту
<span class="muted"> (если Generate ошибся — опишите, что исправить, и нажмите Generate снова)</span> <span class="muted"> (если Generate ошибся — опишите, что исправить, и нажмите Generate снова)</span>
@@ -273,13 +524,36 @@ onMounted(() => {
placeholder="например: description — абзацы между заголовком и координатами, без #хештегов" placeholder="например: description — абзацы между заголовком и координатами, без #хештегов"
/> />
</label> </label>
</template>
<template v-else>
<label>
Instruction
<textarea
v-model="llmInstruction"
class="hint-area"
rows="3"
placeholder="Инструкция для DeepSeek на каждый пост"
/>
</label>
<label class="required-item multi-flag">
<input v-model="llmMultiEvent" type="checkbox" />
Несколько событий в посте
<span class="muted"> (LLM вернёт массив; ingest с URL #e1, #e2…)</span>
</label>
</template>
<div class="form-actions"> <div class="form-actions">
<button <button
type="button" type="button"
class="btn btn-primary" class="btn btn-primary"
:disabled="generating || !samplePost.trim()" :disabled="generating || (profileKind === 'heuristic' && !samplePost.trim())"
@click="handleGenerate" @click="handleGenerate"
> >
<template v-if="profileKind === 'llm'">
{{ generating ? "…" : "Подставить дефолты schema" }}
</template>
<template v-else>
{{ {{
generating generating
? "Генерация..." ? "Генерация..."
@@ -287,11 +561,15 @@ onMounted(() => {
? "Generate / Refine" ? "Generate / Refine"
: "Generate" : "Generate"
}} }}
</template>
</button> </button>
<button <button
type="button" type="button"
class="btn" class="btn"
:disabled="previewing || !profileJson.trim()" :disabled="
previewing ||
(profileKind === 'heuristic' ? !profileJson.trim() : !llmSchemaJson.trim())
"
@click="handlePreview" @click="handlePreview"
> >
{{ previewing ? "Preview..." : "Preview" }} {{ previewing ? "Preview..." : "Preview" }}
@@ -299,7 +577,7 @@ onMounted(() => {
<button <button
type="button" type="button"
class="btn btn-primary" class="btn btn-primary"
:disabled="submitting || !profileJson.trim() || !profileName.trim()" :disabled="submitting || !canSave()"
@click="handleSave" @click="handleSave"
> >
{{ submitting ? "Сохранение..." : editingId != null ? "Обновить" : "Сохранить профиль" }} {{ submitting ? "Сохранение..." : editingId != null ? "Обновить" : "Сохранить профиль" }}
@@ -332,43 +610,120 @@ onMounted(() => {
</section> </section>
<section class="card"> <section class="card">
<h3>Профиль (JSON)</h3> <h3 v-if="profileKind === 'heuristic'">Профиль (JSON)</h3>
<h3 v-else>extract_schema (JSON)</h3>
<textarea <textarea
v-if="profileKind === 'heuristic'"
v-model="profileJson" v-model="profileJson"
class="profile-area" class="profile-area"
rows="16" rows="16"
spellcheck="false" spellcheck="false"
placeholder="Появится после Generate" placeholder="Появится после Generate"
/> />
<textarea
v-else
v-model="llmSchemaJson"
class="profile-area"
rows="16"
spellcheck="false"
placeholder="Ключ → описание поля для LLM"
/>
</section> </section>
</div> </div>
<section v-if="previewFields" class="card"> <section v-if="previewFields" class="card">
<h3>Preview строки</h3> <h3>
Preview
<span v-if="previewEvents.length > 1" class="muted">
({{ previewEvents.length }} событий)
</span>
<span v-else>строки</span>
</h3>
<p v-if="profileKind === 'llm' && !previewIsEvent" class="form-warn">
is_event=false — пост не будет ингеститься.
</p>
<p v-if="profileKind === 'llm' && previewEvents.length > 1" class="form-success match-ok">
Из поста извлечено {{ previewEvents.length }} событий;
пройдут required_fields: {{ previewMatchedCount }}.
Runtime URL: пост#e1 … #e{{ previewEvents.length }}.
</p>
<p v-if="emptyFields.length" class="form-warn"> <p v-if="emptyFields.length" class="form-warn">
Пустые поля: <code>{{ emptyFields.join(", ") }}</code> — уточните подсказку и Generate / Refine. Пустые поля (первая строка): <code>{{ emptyFields.join(", ") }}</code>
<template v-if="profileKind === 'heuristic'">
— уточните подсказку и Generate / Refine.
</template>
</p>
<p v-if="!requiredFields.length" class="form-warn">
Нет обязательных полей —
<template v-if="profileKind === 'llm'">
фильтр только по is_event.
</template>
<template v-else>
любой пост канала будет сохранён. Отметьте поля, без которых пост нужно отбрасывать.
</template>
</p>
<p v-else-if="!previewMatched && previewEvents.length <= 1" class="form-warn">
Этот образец не прошёл бы фильтр. Не хватает:
<code>{{ previewMissing.join(", ") }}</code>
</p>
<p v-else-if="previewMatched" class="form-success match-ok">
<template v-if="previewEvents.length <= 1">
Образец проходит фильтр. Runtime сохранит только посты с заполненными
<code>{{ requiredFields.join(", ") }}</code>.
</template>
<template v-else>
Обязательные поля: <code>{{ requiredFields.join(", ") }}</code>
(проверка на каждое событие).
</template>
</p> </p>
<div class="table-wrap"> <div class="table-wrap">
<table class="admin-table"> <table class="admin-table">
<thead> <thead>
<tr> <tr>
<th v-if="previewEvents.length > 1">#</th>
<th v-for="key in Object.keys(previewFields)" :key="key">{{ key }}</th> <th v-for="key in Object.keys(previewFields)" :key="key">{{ key }}</th>
</tr> </tr>
</thead> </thead>
<tbody> <tbody>
<tr> <tr v-for="(row, idx) in (previewEvents.length ? previewEvents : [previewFields])" :key="idx">
<td v-if="previewEvents.length > 1">{{ idx + 1 }}</td>
<td <td
v-for="(val, key) in previewFields" v-for="key in Object.keys(previewFields)"
:key="key" :key="key"
class="preview-cell" class="preview-cell"
:class="{ empty: !val }" :class="{ empty: !row?.[key] }"
> >
{{ val || "—" }} {{ row?.[key] || "—" }}
</td> </td>
</tr> </tr>
</tbody> </tbody>
</table> </table>
</div> </div>
<div class="required-box">
<h4>Обязательно для сохранения</h4>
<p class="muted required-hint">
<template v-if="profileKind === 'llm'">
После LLM: событие отбрасывается, если обязательное поле пустое (поверх is_event).
</template>
<template v-else>
Пост отбрасывается, если хоть одно отмеченное поле не извлеклось (для coords — ещё и не парсится lat,lon).
</template>
</p>
<div class="required-grid">
<label
v-for="key in Object.keys(previewFields)"
:key="key"
class="required-item"
>
<input
type="checkbox"
:checked="isRequired(key)"
@change="toggleRequired(key)"
/>
<code>{{ key }}</code>
</label>
</div>
</div>
</section> </section>
<p v-if="error" class="form-error">{{ error }}</p> <p v-if="error" class="form-error">{{ error }}</p>
@@ -382,6 +737,7 @@ onMounted(() => {
<tr> <tr>
<th>ID</th> <th>ID</th>
<th>Название</th> <th>Название</th>
<th>Тип</th>
<th>Статус</th> <th>Статус</th>
<th>Создан</th> <th>Создан</th>
<th></th> <th></th>
@@ -391,6 +747,7 @@ onMounted(() => {
<tr v-for="p in profiles" :key="p.id"> <tr v-for="p in profiles" :key="p.id">
<td>{{ p.id }}</td> <td>{{ p.id }}</td>
<td>{{ p.name }}</td> <td>{{ p.name }}</td>
<td><code>{{ kindLabel(p.kind) }}</code></td>
<td>{{ p.status }}</td> <td>{{ p.status }}</td>
<td>{{ formatDate(p.created_at) }}</td> <td>{{ formatDate(p.created_at) }}</td>
<td class="actions"> <td class="actions">
@@ -401,7 +758,7 @@ onMounted(() => {
</td> </td>
</tr> </tr>
<tr v-if="!loading && !profiles.length"> <tr v-if="!loading && !profiles.length">
<td colspan="5" class="muted">Пока нет профилей</td> <td colspan="6" class="muted">Пока нет профилей</td>
</tr> </tr>
</tbody> </tbody>
</table> </table>
@@ -486,6 +843,42 @@ onMounted(() => {
margin: 0 0 0.75rem; margin: 0 0 0.75rem;
} }
.match-ok {
margin: 0 0 0.75rem;
}
.required-box {
margin-top: 1rem;
}
.required-box h4 {
margin: 0 0 0.25rem;
font-size: 0.9rem;
}
.required-hint {
margin: 0 0 0.6rem;
font-size: 0.8rem;
}
.required-grid {
display: flex;
flex-wrap: wrap;
gap: 0.5rem 1rem;
}
.required-item {
display: flex;
align-items: center;
gap: 0.35rem;
font-size: 0.85rem;
}
.multi-flag {
margin: 0.75rem 0 0;
flex-wrap: wrap;
}
.actions { .actions {
display: flex; display: flex;
gap: 0.35rem; gap: 0.35rem;
@@ -38,7 +38,11 @@ const editForm = ref({
}); });
const readyProfiles = computed(() => const readyProfiles = computed(() =>
profiles.value.filter((p) => !!p.heuristic_profile), profiles.value.filter((p) => {
if (p.status !== "ready") return false;
if (p.kind === "llm") return !!p.llm_profile;
return !!p.heuristic_profile;
}),
); );
const activeChannels = computed(() => channels.value.filter((c) => c.is_active)); const activeChannels = computed(() => channels.value.filter((c) => c.is_active));
@@ -186,8 +190,17 @@ async function handleDelete(job: ParseJob) {
} }
} }
const STALE_JOB_MS = 15 * 60 * 1000;
function isStaleJob(job: ParseJob): boolean {
if (job.status !== "queued" && job.status !== "running") return false;
if (!job.last_run_at) return job.status === "running";
return Date.now() - new Date(job.last_run_at).getTime() >= STALE_JOB_MS;
}
function canForceRun(job: ParseJob): boolean { function canForceRun(job: ParseJob): boolean {
return job.status !== "queued" && job.status !== "running"; if (job.status !== "queued" && job.status !== "running") return true;
return isStaleJob(job);
} }
async function handleForceRun(jobId: number) { async function handleForceRun(jobId: number) {
@@ -237,9 +250,10 @@ onUnmounted(() => {
Создайте Создайте
<router-link to="/channels">канал</router-link> <router-link to="/channels">канал</router-link>
и и
<router-link to="/parser-profiles">профиль</router-link>, <router-link to="/parser-profiles">профиль</router-link>
затем свяжите их здесь. В Redis уходит плоский (heuristic или llm), затем свяжите их здесь. Flatten в Redis:
<code>extract_mode=profile</code> + <code>heuristic_profile</code>. heuristic → <code>extract_mode=profile</code>;
llm → <code>extract_mode=llm</code> (batch; listener для llm пока fallback).
</p> </p>
<form class="admin-form" @submit.prevent="handleCreatePair"> <form class="admin-form" @submit.prevent="handleCreatePair">
<div class="form-row"> <div class="form-row">
@@ -257,7 +271,7 @@ onUnmounted(() => {
<select v-model.number="pairForm.profile_id" required> <select v-model.number="pairForm.profile_id" required>
<option :value="null" disabled>Выберите…</option> <option :value="null" disabled>Выберите…</option>
<option v-for="p in readyProfiles" :key="p.id" :value="p.id"> <option v-for="p in readyProfiles" :key="p.id" :value="p.id">
{{ p.name }} (#{{ p.id }}) {{ p.name }} (#{{ p.id }}, {{ p.kind === "llm" ? "llm" : "heuristic" }})
</option> </option>
</select> </select>
</label> </label>
@@ -343,7 +357,7 @@ onUnmounted(() => {
<button <button
class="btn btn-sm btn-danger" class="btn btn-sm btn-danger"
type="button" type="button"
:disabled="job.status === 'running'" :disabled="job.status === 'running' && !isStaleJob(job)"
@click="handleDelete(job)" @click="handleDelete(job)"
> >
Удалить Удалить
+28 -6
View File
@@ -54,14 +54,29 @@ async def run(job_id, source_config, *, ctx) -> tuple[list[dict], str | None]
- `source_type` события = тип адаптера; - `source_type` события = тип адаптера;
- `source_config` валидируется схемами из `contracts/sources.py`. - `source_config` валидируется схемами из `contracts/sources.py`.
### Режимы извлечения (Telegram)
Цель всегда одна: фиксированные поля `IngestEventItem` / Event в ЦА (`title`, `description`, `locality`, `coords`→lat/lng, `event_date`, `topic`, …). Кастомные пользовательские таблицы — roadmap.
| `extract_mode` | Что делает | Откуда в UI |
|----------------|------------|-------------|
| `heuristic` | Legacy regex-парсер постов | raw job API (не пара) |
| `llm` | DeepSeek **на каждый** пост (batch) | профиль `kind=llm` → pair flatten |
| `profile` | Статичные правила `HeuristicProfile` | профиль `kind=heuristic` → pair flatten |
### LLM-режим (`extract_mode: llm`) ### LLM-режим (`extract_mode: llm`)
Для неструктурированных Telegram-постов и Crawl4AI: Для неструктурированных Telegram-постов и Crawl4AI:
- ключ `DEEPSEEK_API_KEY` в `.env` (воркеры `cp-workers` / `cp-workers-web`); - ключ `DEEPSEEK_API_KEY` в `.env` (воркеры `cp-workers` / `cp-workers-web`; preview в `ca-api`);
- Telegram: текст поста → DeepSeek JSON → `IngestEventItem`; - Telegram batch: текст → DeepSeek JSON → один или несколько `IngestEventItem` (`workers/llm_extract.py`);
- опционально `extract_schema`, `instruction`, `required_fields` (post-extract gate поверх `is_event`);
- **`multi_event`** (opt-in в `llm_profile` / `TelegramSourceConfig`): модель возвращает массив `events`; каждое событие — отдельный ingest. URL: один event → `post.url`; несколько → `post.url#e1`, `#e2`, … (дедуп ЦА по `source_url`);
- heuristic / profile по-прежнему **1 пост → 1 событие**;
- Crawl4AI: страница → `LLMExtractionStrategy` (DeepSeek) с fallback на тот же DeepSeek по markdown; - Crawl4AI: страница → `LLMExtractionStrategy` (DeepSeek) с fallback на тот же DeepSeek по markdown;
- в UI «Парсеры»: поле **Извлечение** = LLM DeepSeek. - **listener:** для `llm` пока fallback на heuristic (LLM — batch-only by design).
Reusable LLM-профиль в ЦА: `ParserProfile.kind=llm` + JSON `llm_profile` (`contracts/llm_profile.py`). При enqueue flatten → `extract_mode=llm` + schema/instruction/required_fields/`multi_event`.
### Profile-режим (`extract_mode: profile`) ### Profile-режим (`extract_mode: profile`)
@@ -69,8 +84,12 @@ async def run(job_id, source_config, *, ctx) -> tuple[list[dict], str | None]
- `heuristic_profile` в `source_config` (схема `contracts/heuristic_profile.py`); - `heuristic_profile` в `source_config` (схема `contracts/heuristic_profile.py`);
- интерпретатор: `workers/heuristic_profile.py` (те же правила, что preview в ЦА); - интерпретатор: `workers/heuristic_profile.py` (те же правила, что preview в ЦА);
- Telegram batch + listener применяют профиль без DeepSeek; - `required_fields`: пост без заполненных обязательных полей не ингестится (batch + listener);
- генерация профиля — только в админке (`/admin/parser-profiles/generate`). - генерация правил — один раз в админке (`/admin/parser-profiles/generate`); runtime **без** LLM.
Reusable heuristic-профиль: `ParserProfile.kind=heuristic` + `heuristic_profile`. Pair на «Парсеры» → flatten `extract_mode=profile`.
UI «Профили»: выбор `kind` (heuristic | llm). Связка канал+профиль на «Парсеры».
```mermaid ```mermaid
flowchart LR flowchart LR
@@ -127,7 +146,10 @@ Legacy-ключ `cp:jobs` по-прежнему дренируется telegram-
## Telegram real-time ## Telegram real-time
`TelegramListener` без изменений: подписки из ЦА, ingest с `listener: true` (статус `ParseJob` не трогается). `TelegramListener`: подписки из ЦА, ingest с `listener: true` (статус `ParseJob` не трогается).
- `extract_mode=profile` — те же heuristic-правила, что batch;
- `extract_mode=llm` — **не** вызывает DeepSeek; fallback на legacy heuristic (LLM только в batch).
## Переменные окружения ## Переменные окружения
@@ -11,9 +11,11 @@ from workers.heuristic_profile import extract_with_profile
from workers.llm_extract import ( from workers.llm_extract import (
DEFAULT_EXTRACT_SCHEMA, DEFAULT_EXTRACT_SCHEMA,
DEFAULT_INSTRUCTION, DEFAULT_INSTRUCTION,
extract_event_fields, event_source_url,
extract_event_list,
fields_to_ingest, fields_to_ingest,
llm_enabled, llm_enabled,
match_llm_required,
) )
from workers.parsers.telegram_events import parse_event_posts from workers.parsers.telegram_events import parse_event_posts
from workers.sources.telegram_client import ( from workers.sources.telegram_client import (
@@ -73,8 +75,7 @@ def _extract_posts_with_profile(posts, cfg: TelegramSourceConfig) -> tuple[list[
text = (post.text or "").strip() text = (post.text or "").strip()
if not text: if not text:
continue continue
events.append( event = extract_with_profile(
extract_with_profile(
text, text,
profile, profile,
source_url=post.url, source_url=post.url,
@@ -85,32 +86,48 @@ def _extract_posts_with_profile(posts, cfg: TelegramSourceConfig) -> tuple[list[
"post_date": post.date.isoformat() if post.date else None, "post_date": post.date.isoformat() if post.date else None,
}, },
) )
) if event is None:
logger.debug("Profile skip %s (required fields missing)", post.url)
continue
events.append(event)
return events, None return events, None
async def _extract_posts_with_llm(posts, cfg: TelegramSourceConfig) -> tuple[list[dict], str | None]: async def _extract_posts_with_llm(posts, cfg: TelegramSourceConfig) -> tuple[list[dict], str | None]:
schema = cfg.extract_schema or DEFAULT_EXTRACT_SCHEMA schema = cfg.extract_schema or DEFAULT_EXTRACT_SCHEMA
instruction = cfg.instruction or DEFAULT_INSTRUCTION instruction = cfg.instruction or DEFAULT_INSTRUCTION
multi_event = bool(cfg.multi_event)
events: list[dict] = [] events: list[dict] = []
errors: list[str] = [] errors: list[str] = []
required = list(cfg.required_fields or [])
for post in posts: for post in posts:
text = (post.text or "").strip() text = (post.text or "").strip()
if not text: if not text:
continue continue
try: try:
fields = await extract_event_fields( field_list = await extract_event_list(
text, text,
extract_schema=schema, extract_schema=schema,
instruction=instruction, instruction=instruction,
multi_event=multi_event,
)
if not field_list:
continue
total = len(field_list)
for idx, fields in enumerate(field_list):
if not match_llm_required(fields, required):
logger.debug(
"LLM skip %s event %s/%s (required fields missing)",
post.url,
idx + 1,
total,
) )
if not fields.get("is_event", True):
continue continue
events.append( events.append(
fields_to_ingest( fields_to_ingest(
source_type="telegram", source_type="telegram",
source_url=post.url, source_url=event_source_url(post.url, idx, total),
raw_text=text, raw_text=text,
fields=fields, fields=fields,
domain_profile="telegram_llm", domain_profile="telegram_llm",
@@ -118,6 +135,9 @@ async def _extract_posts_with_llm(posts, cfg: TelegramSourceConfig) -> tuple[lis
"channel": post.channel, "channel": post.channel,
"message_id": post.id, "message_id": post.id,
"post_date": post.date.isoformat() if post.date else None, "post_date": post.date.isoformat() if post.date else None,
"event_index": idx + 1,
"event_count": total,
"multi_event": multi_event and total > 1,
}, },
) )
) )
@@ -4,7 +4,7 @@ from __future__ import annotations
from typing import Any from typing import Any
from contracts.heuristic_profile import HeuristicProfile, apply_profile from contracts.heuristic_profile import HeuristicProfile, match_profile
from workers.llm_extract import parse_coords, parse_date from workers.llm_extract import parse_coords, parse_date
@@ -62,8 +62,10 @@ def extract_with_profile(
source_url: str, source_url: str,
source_type: str = "telegram", source_type: str = "telegram",
extra_metadata: dict | None = None, extra_metadata: dict | None = None,
) -> dict: ) -> dict | None:
fields = apply_profile(text, profile) matched, fields, _missing = match_profile(text, profile)
if not matched:
return None
return profile_fields_to_ingest( return profile_fields_to_ingest(
source_type=source_type, source_type=source_type,
source_url=source_url, source_url=source_url,
+128 -48
View File
@@ -5,30 +5,31 @@ from __future__ import annotations
import json import json
import logging import logging
import os import os
import re
from datetime import datetime, timezone from datetime import datetime, timezone
from typing import Any from typing import Any
import httpx import httpx
from contracts.heuristic_profile import parse_coords
from contracts.llm_profile import (
DEFAULT_EXTRACT_SCHEMA,
DEFAULT_INSTRUCTION,
match_llm_required,
)
logger = logging.getLogger("cp-worker.llm") logger = logging.getLogger("cp-worker.llm")
COORDS_RE = re.compile(r"(-?\d{1,3}\.\d+)\s*,\s*(-?\d{1,3}\.\d+)") # Re-export for adapters that import from this module
__all__ = [
DEFAULT_EXTRACT_SCHEMA: dict[str, str] = { "DEFAULT_EXTRACT_SCHEMA",
"title": "string — short event title", "DEFAULT_INSTRUCTION",
"locality": "string — place / settlement name", "event_source_url",
"event_date": "string — date as DD.MM.YYYY or YYYY-MM-DD if known", "extract_event_fields",
"description": "string — concise event summary", "extract_event_list",
"coords": "string — latitude, longitude if present else empty", "fields_to_ingest",
"topic": "string — short topic tag", "llm_enabled",
} "match_llm_required",
]
DEFAULT_INSTRUCTION = (
"Extract structured military/news event fields from the text. "
"If the text is not an event, return is_event=false. "
"Respond with a single JSON object only."
)
def llm_enabled() -> bool: def llm_enabled() -> bool:
@@ -43,29 +44,57 @@ def llm_settings() -> dict[str, str]:
} }
async def extract_event_fields( def event_source_url(post_url: str, index: int, total: int) -> str:
text: str, """Stable per-event URL for CA dedup. Single event keeps bare post URL."""
base = (post_url or "").strip()
if total <= 1:
return base
# index is 0-based; fragment uses 1-based #eN
return f"{base}#e{index + 1}"
def _normalize_fields(raw: dict[str, Any], schema: dict[str, str]) -> dict[str, str]:
return {key: str(raw.get(key) or "").strip() for key in schema}
def _parse_llm_content(
content: str,
schema: dict[str, str],
*, *,
extract_schema: dict[str, str] | None = None, multi_event: bool,
instruction: str | None = None, ) -> tuple[bool, list[dict[str, str]]]:
) -> dict[str, Any]: """Return (is_event, list of field dicts)."""
"""Ask DeepSeek to fill schema fields from free text. Returns dict (+ is_event).""" parsed = json.loads(content)
if not isinstance(parsed, dict):
return False, []
is_event = bool(parsed.get("is_event", True))
if not is_event:
return False, []
if multi_event and isinstance(parsed.get("events"), list):
events: list[dict[str, str]] = []
for item in parsed["events"]:
if isinstance(item, dict):
events.append(_normalize_fields(item, schema))
return True, events
# Single-event fallback (legacy fields / top-level keys)
fields_raw = parsed.get("fields") if isinstance(parsed.get("fields"), dict) else parsed
if not isinstance(fields_raw, dict):
fields_raw = {}
# Drop non-schema keys that confuse normalize when falling back to top-level
cleaned = {k: fields_raw.get(k) for k in schema}
return True, [_normalize_fields(cleaned, schema)]
async def _call_deepseek(user_prompt: str) -> str:
settings = llm_settings() settings = llm_settings()
if not settings["api_key"]: if not settings["api_key"]:
raise RuntimeError( raise RuntimeError(
"DEEPSEEK_API_KEY is not set. Add it to .env for LLM extract_mode." "DEEPSEEK_API_KEY is not set. Add it to .env for LLM extract_mode."
) )
schema = extract_schema or DEFAULT_EXTRACT_SCHEMA
instr = instruction or DEFAULT_INSTRUCTION
schema_lines = "\n".join(f"- {k}: {v}" for k, v in schema.items())
user_prompt = (
f"{instr}\n\n"
f"Fields to extract:\n{schema_lines}\n\n"
'Return JSON: {"is_event": true|false, "fields": {<field>: <string>}}\n\n'
f"Text:\n{text[:12000]}"
)
payload = { payload = {
"model": settings["model"], "model": settings["model"],
"messages": [ "messages": [
@@ -95,14 +124,72 @@ async def extract_event_fields(
response.raise_for_status() response.raise_for_status()
data = response.json() data = response.json()
content = data["choices"][0]["message"]["content"] return data["choices"][0]["message"]["content"]
parsed = json.loads(content)
fields = parsed.get("fields") if isinstance(parsed.get("fields"), dict) else parsed
if not isinstance(fields, dict): async def extract_event_list(
fields = {} text: str,
# Normalize to strings for known keys *,
result = {key: str(fields.get(key) or "").strip() for key in schema} extract_schema: dict[str, str] | None = None,
result["is_event"] = bool(parsed.get("is_event", True)) instruction: str | None = None,
multi_event: bool = False,
) -> list[dict[str, Any]]:
"""Extract zero or more event field dicts from free text.
Each dict has schema keys as strings. Empty list if not an event / no items.
"""
schema = extract_schema or DEFAULT_EXTRACT_SCHEMA
instr = instruction or DEFAULT_INSTRUCTION
schema_lines = "\n".join(f"- {k}: {v}" for k, v in schema.items())
if multi_event:
user_prompt = (
f"{instr}\n\n"
"If the text describes multiple distinct events (different places, "
"coords, or dates), return one object per event in \"events\".\n"
f"Fields per event:\n{schema_lines}\n\n"
'Return JSON: {"is_event": true|false, "events": [{<field>: <string>}, ...]}\n'
"If there is no event, return is_event=false and events=[].\n\n"
f"Text:\n{text[:12000]}"
)
else:
user_prompt = (
f"{instr}\n\n"
f"Fields to extract:\n{schema_lines}\n\n"
'Return JSON: {"is_event": true|false, "fields": {<field>: <string>}}\n\n'
f"Text:\n{text[:12000]}"
)
content = await _call_deepseek(user_prompt)
is_event, events = _parse_llm_content(content, schema, multi_event=multi_event)
if not is_event:
return []
return events
async def extract_event_fields(
text: str,
*,
extract_schema: dict[str, str] | None = None,
instruction: str | None = None,
) -> dict[str, Any]:
"""Ask DeepSeek to fill schema fields from free text. Returns dict (+ is_event).
Single-event API kept for crawl4ai and legacy callers.
"""
events = await extract_event_list(
text,
extract_schema=extract_schema,
instruction=instruction,
multi_event=False,
)
if not events:
schema = extract_schema or DEFAULT_EXTRACT_SCHEMA
empty = {key: "" for key in schema}
empty["is_event"] = False
return empty
result = dict(events[0])
result["is_event"] = True
return result return result
@@ -146,13 +233,6 @@ def fields_to_ingest(
} }
def parse_coords(raw: str) -> tuple[float | None, float | None]:
match = COORDS_RE.search(raw or "")
if not match:
return None, None
return float(match.group(1)), float(match.group(2))
def parse_date(raw: str) -> datetime | None: def parse_date(raw: str) -> datetime | None:
if not raw: if not raw:
return None return None
@@ -81,7 +81,7 @@ class TelegramListener:
", ".join(sorted(channels)) or "(none)", ", ".join(sorted(channels)) or "(none)",
) )
def _build_event(self, channel: str, post) -> dict: def _build_event(self, channel: str, post) -> dict | None:
sub = self._channels.get(normalize_channel(channel)) or {} sub = self._channels.get(normalize_channel(channel)) or {}
cfg = sub.get("source_config") or {} cfg = sub.get("source_config") or {}
extract_mode = cfg.get("extract_mode") or "heuristic" extract_mode = cfg.get("extract_mode") or "heuristic"
@@ -107,6 +107,9 @@ class TelegramListener:
async def _ingest_post(self, channel: str, post) -> None: async def _ingest_post(self, channel: str, post) -> None:
event = self._build_event(channel, post) event = self._build_event(channel, post)
if event is None:
logger.debug("Listener skip %s (profile required fields missing)", post.url)
return
sub = self._channels.get(normalize_channel(channel)) or {} sub = self._channels.get(normalize_channel(channel)) or {}
job_id = sub.get("job_id") job_id = sub.get("job_id")
+53
View File
@@ -17,6 +17,8 @@ TARGET_FIELDS: tuple[str, ...] = (
"region", "region",
) )
COORDS_RE = re.compile(r"(-?\d{1,3}\.\d+)\s*,\s*(-?\d{1,3}\.\d+)")
Strategy = Literal["regex", "line", "after_marker", "between", "full_text", "literal"] Strategy = Literal["regex", "line", "after_marker", "between", "full_text", "literal"]
@@ -45,8 +47,24 @@ class FieldRule(BaseModel):
class HeuristicProfile(BaseModel): class HeuristicProfile(BaseModel):
version: Literal[1] = 1 version: Literal[1] = 1
fields: dict[str, FieldRule] = Field(default_factory=dict) fields: dict[str, FieldRule] = Field(default_factory=dict)
required_fields: list[str] = Field(default_factory=list)
notes: str = "" notes: str = ""
@field_validator("required_fields")
@classmethod
def known_required(cls, value: list[str]) -> list[str]:
seen: list[str] = []
unknown: list[str] = []
for name in value:
if name not in TARGET_FIELDS:
unknown.append(name)
continue
if name not in seen:
seen.append(name)
if unknown:
raise ValueError(f"Unknown required_fields: {sorted(set(unknown))}")
return seen
@model_validator(mode="after") @model_validator(mode="after")
def known_fields_only(self) -> "HeuristicProfile": def known_fields_only(self) -> "HeuristicProfile":
unknown = set(self.fields) - set(TARGET_FIELDS) unknown = set(self.fields) - set(TARGET_FIELDS)
@@ -163,3 +181,38 @@ def apply_profile(text: str, profile: HeuristicProfile | dict[str, Any]) -> dict
for name, rule in profile.fields.items(): for name, rule in profile.fields.items():
result[name] = _apply_rule(text, rule) result[name] = _apply_rule(text, rule)
return result return result
def parse_coords(raw: str) -> tuple[float | None, float | None]:
match = COORDS_RE.search(raw or "")
if not match:
return None, None
return float(match.group(1)), float(match.group(2))
def _field_is_filled(name: str, value: str) -> bool:
if not (value or "").strip():
return False
if name == "coords":
lat, lng = parse_coords(value)
return lat is not None and lng is not None
return True
def match_profile(
text: str,
profile: HeuristicProfile | dict[str, Any],
) -> tuple[bool, dict[str, str], list[str]]:
"""Apply profile and report whether required_fields are filled.
Empty required_fields means no gate (legacy profiles ingest every post).
"""
if isinstance(profile, dict):
profile = HeuristicProfile.model_validate(profile)
fields = apply_profile(text, profile)
missing = [
name
for name in profile.required_fields
if not _field_is_filled(name, fields.get(name, ""))
]
return (not missing, fields, missing)
+86
View File
@@ -0,0 +1,86 @@
"""Reusable LLM extract profile (CA storage + flattened into TelegramSourceConfig)."""
from __future__ import annotations
from typing import Any
from pydantic import BaseModel, Field, field_validator, model_validator
from contracts.heuristic_profile import TARGET_FIELDS
# LLM defaults omit region (same as workers/llm_extract.DEFAULT_EXTRACT_SCHEMA).
LLM_SCHEMA_FIELDS: tuple[str, ...] = (
"title",
"locality",
"event_date",
"description",
"coords",
"topic",
)
DEFAULT_EXTRACT_SCHEMA: dict[str, str] = {
"title": "string — short event title",
"locality": "string — place / settlement name",
"event_date": "string — date as DD.MM.YYYY or YYYY-MM-DD if known",
"description": "string — concise event summary",
"coords": "string — latitude, longitude if present else empty",
"topic": "string — short topic tag",
}
DEFAULT_INSTRUCTION = (
"Extract structured military/news event fields from the text. "
"If the text is not an event, return is_event=false. "
"Respond with a single JSON object only."
)
class LlmProfile(BaseModel):
instruction: str | None = None
extract_schema: dict[str, str] = Field(default_factory=lambda: dict(DEFAULT_EXTRACT_SCHEMA))
required_fields: list[str] = Field(default_factory=list)
# When true, LLM may return multiple events per post (array "events")
multi_event: bool = False
@field_validator("instruction", mode="before")
@classmethod
def empty_instruction_to_none(cls, value: Any) -> Any:
if value is None:
return None
if isinstance(value, str) and not value.strip():
return None
return value
@field_validator("required_fields")
@classmethod
def known_required(cls, value: list[str]) -> list[str]:
seen: list[str] = []
unknown: list[str] = []
for name in value:
if name not in TARGET_FIELDS:
unknown.append(name)
continue
if name not in seen:
seen.append(name)
if unknown:
raise ValueError(f"Unknown required_fields: {sorted(set(unknown))}")
return seen
@model_validator(mode="after")
def known_schema_keys(self) -> "LlmProfile":
if not self.extract_schema:
raise ValueError("extract_schema must not be empty")
unknown = set(self.extract_schema) - set(TARGET_FIELDS)
if unknown:
raise ValueError(f"Unknown extract_schema fields: {sorted(unknown)}")
return self
def match_llm_required(fields: dict[str, Any], required_fields: list[str]) -> bool:
"""True if all required_fields are non-empty strings (empty required = no gate)."""
if not required_fields:
return True
for name in required_fields:
val = fields.get(name)
if val is None or not str(val).strip():
return False
return True
+29 -3
View File
@@ -15,6 +15,10 @@ class TelegramSourceConfig(BaseModel):
extract_schema: dict[str, str] | None = None extract_schema: dict[str, str] | None = None
instruction: str | None = None instruction: str | None = None
heuristic_profile: dict | None = None heuristic_profile: dict | None = None
# Post-extract gate for extract_mode=llm (empty = only is_event filter)
required_fields: list[str] = Field(default_factory=list)
# Split one post into N events when extract_mode=llm (opt-in)
multi_event: bool = False
sample_post: str | None = None # audit / re-generate; not required at runtime sample_post: str | None = None # audit / re-generate; not required at runtime
@field_validator("channel") @field_validator("channel")
@@ -22,15 +26,37 @@ class TelegramSourceConfig(BaseModel):
def strip_channel(cls, value: str) -> str: def strip_channel(cls, value: str) -> str:
return value.strip() return value.strip()
@field_validator("required_fields")
@classmethod
def known_required(cls, value: list[str]) -> list[str]:
from contracts.heuristic_profile import TARGET_FIELDS
seen: list[str] = []
unknown: list[str] = []
for name in value:
if name not in TARGET_FIELDS:
unknown.append(name)
continue
if name not in seen:
seen.append(name)
if unknown:
raise ValueError(f"Unknown required_fields: {sorted(set(unknown))}")
return seen
@model_validator(mode="after") @model_validator(mode="after")
def profile_requires_rules(self) -> "TelegramSourceConfig": def validate_extract_mode(self) -> "TelegramSourceConfig":
if self.extract_mode != "profile": if self.extract_mode == "profile":
return self
if not self.heuristic_profile: if not self.heuristic_profile:
raise ValueError("heuristic_profile required when extract_mode=profile") raise ValueError("heuristic_profile required when extract_mode=profile")
from contracts.heuristic_profile import HeuristicProfile from contracts.heuristic_profile import HeuristicProfile
HeuristicProfile.model_validate(self.heuristic_profile) HeuristicProfile.model_validate(self.heuristic_profile)
elif self.extract_mode == "llm" and self.extract_schema is not None:
from contracts.heuristic_profile import TARGET_FIELDS
unknown = set(self.extract_schema) - set(TARGET_FIELDS)
if unknown:
raise ValueError(f"Unknown extract_schema fields: {sorted(unknown)}")
return self return self
+7 -1
View File
@@ -184,7 +184,12 @@ async def run(job_id, source_config, *, ctx) -> tuple[list[dict], str | None]
Опционально: `extract_mode: llm` (DeepSeek) для telegram/crawl4ai в batch. Опционально: `extract_mode: llm` (DeepSeek) для telegram/crawl4ai в batch.
**Профиль + канал:** UI `/parser-profiles` (Generate/Preview/CRUD) и `/channels`; связка на `/parsers` → `ParseJob` с FK. При enqueue ЦА flatten в `extract_mode=profile` + `heuristic_profile`. Runtime (batch + listener) применяет только профиль, без LLM. Кастомные пользовательские таблицы — roadmap. **Профиль + канал:** UI `/parser-profiles` (CRUD, Generate/Preview для heuristic; instruction/schema/Preview для llm) и `/channels`; связка на `/parsers` → `ParseJob` с FK. При enqueue ЦА flatten по `ParserProfile.kind`:
- `heuristic` → `extract_mode=profile` + `heuristic_profile` (runtime без LLM; listener поддерживает);
- `llm` → `extract_mode=llm` + `extract_schema` / `instruction` / `required_fields` (DeepSeek на каждый пост в batch; listener пока fallback на heuristic).
Целевые поля — фиксированная схема Event/`IngestEventItem`. Кастомные пользовательские таблицы — roadmap.
Подробности: [centers/parsing/ARCHITECTURE.md](../centers/parsing/ARCHITECTURE.md), [centers/analytics/ARCHITECTURE.md](../centers/analytics/ARCHITECTURE.md). Подробности: [centers/parsing/ARCHITECTURE.md](../centers/parsing/ARCHITECTURE.md), [centers/analytics/ARCHITECTURE.md](../centers/analytics/ARCHITECTURE.md).
@@ -209,6 +214,7 @@ async def run(job_id, source_config, *, ctx) -> tuple[list[dict], str | None]
| `contracts/sources.py` | Валидация `source_config` по `source_type` | | `contracts/sources.py` | Валидация `source_config` по `source_type` |
| `contracts/queues.py` | `SOURCE_FAMILY` → ключ очереди | | `contracts/queues.py` | `SOURCE_FAMILY` → ключ очереди |
| `contracts/heuristic_profile.py` | Статичный профиль конструктора + `apply_profile` | | `contracts/heuristic_profile.py` | Статичный профиль конструктора + `apply_profile` |
| `contracts/llm_profile.py` | Reusable LLM instruction/schema + `match_llm_required` |
Правило: меняете форму события / конфиг источника / очередь — сначала `contracts/`, потом CA/CP/UI. Правило: меняете форму события / конфиг источника / очередь — сначала `contracts/`, потом CA/CP/UI.
+11 -2
View File
@@ -14,6 +14,8 @@
| [`ingest.py`](../contracts/ingest.py) |Контракт результата. Событие и пакет ingest | | [`ingest.py`](../contracts/ingest.py) |Контракт результата. Событие и пакет ingest |
| [`sources.py`](../contracts/sources.py) | Контракт настроек парсера. Схемы `source_config` по `source_type` | | [`sources.py`](../contracts/sources.py) | Контракт настроек парсера. Схемы `source_config` по `source_type` |
| [`queues.py`](../contracts/queues.py) | Контракт доставки задания нужному воркеру. `source_type` → family → Redis key | | [`queues.py`](../contracts/queues.py) | Контракт доставки задания нужному воркеру. `source_type` → family → Redis key |
| [`heuristic_profile.py`](../contracts/heuristic_profile.py) | Статичные правила extract_mode=profile |
| [`llm_profile.py`](../contracts/llm_profile.py) | Reusable LLM instruction/schema + required_fields + `multi_event` |
ЦА при admin CRUD валидирует конфиг через `parse_source_config`. ЦА при admin CRUD валидирует конфиг через `parse_source_config`.
ЦП адаптеры должны отдавать dict, совместимые с `IngestEventItem`. ЦП адаптеры должны отдавать dict, совместимые с `IngestEventItem`.
@@ -57,13 +59,20 @@ Listener добавляет флаг `listener: true` на стороне CA API
| source_type | Модель | Главные поля | | source_type | Модель | Главные поля |
|-------------|--------|--------------| |-------------|--------|--------------|
| `telegram` | `TelegramSourceConfig` | `channel`, `limit`, `extract_mode` (`heuristic`\|`llm`\|`profile`), `heuristic_profile` | | `telegram` | `TelegramSourceConfig` | `channel`, `limit`, `extract_mode` (`heuristic`\|`llm`\|`profile`), `heuristic_profile`, `extract_schema`, `instruction`, `required_fields` (llm gate) |
| `crawl4ai` | `Crawl4AISourceConfig` | `urls`, `extract_mode`, `extract_schema`, `domain_profile` | | `crawl4ai` | `Crawl4AISourceConfig` | `urls`, `extract_mode`, `extract_schema`, `domain_profile` |
| `viina` | `ViinaSourceConfig` | `urls` / `texts`, `input_mode` | | `viina` | `ViinaSourceConfig` | `urls` / `texts`, `input_mode` |
Реестр: `CONFIG_MODELS` + `parse_source_config(source_type, raw)`. Реестр: `CONFIG_MODELS` + `parse_source_config(source_type, raw)`.
Профиль парсера: `contracts/heuristic_profile.py` (`HeuristicProfile`, `apply_profile`). Сущности Profile/Channel живут в БД ЦА; в Redis уходит плоский `heuristic_profile`. Профили в БД ЦА (`ParserProfile`):
| kind | Хранилище | Flatten в Redis |
|------|-----------|-----------------|
| `heuristic` | `heuristic_profile` (`contracts/heuristic_profile.py`) | `extract_mode=profile` |
| `llm` | `llm_profile` (`contracts/llm_profile.py`: instruction, extract_schema, required_fields, multi_event) | `extract_mode=llm` (+ `#eN` URLs if multi) |
`required_fields` — обязательные поля; пустой список = без фильтра (для llm остаётся только `is_event`). Канал + профиль живут в ЦА; CP получает только плоский `source_config`.
--- ---