Add reusable LLM parser profiles with multi-event extract.

Support kind=llm profiles (instruction/schema), optional multi-event posts via #eN URLs, and recover stale running/queued parse jobs after worker crashes.

Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
2026-09-13 20:22:38 +03:00
co-authored by Cursor
parent 5811ecb134
commit 3f9dc6643b
27 changed files with 1334 additions and 205 deletions
+3
View File
@@ -27,6 +27,9 @@ ADMIN_JWT_SECRET=change-me-jwt-secret
# DEEPSEEK_BASE_URL=https://api.deepseek.com
# DEEPSEEK_MODEL=deepseek-chat
# Recover parse jobs stuck in running/queued after worker crash (seconds, default 900)
# STALE_JOB_SECONDS=900
# CP adapter workers (set in docker-compose; override locally if needed)
# ENABLED_ADAPTERS=telegram
# WORKER_FAMILIES=telegram
+1 -1
View File
@@ -74,7 +74,7 @@ Adapters return dicts matching `IngestEventItem` (`contracts/ingest.py`). Requir
### LLM extract (runtime, not dev)
`extract_mode: llm` uses DeepSeek via `workers/llm_extract.py`. Key: `DEEPSEEK_API_KEY` in `.env`. Do not confuse with Cursor dev agents.
`extract_mode: llm` uses DeepSeek via `workers/llm_extract.py` (batch). Reusable profiles in admin UI (`kind=heuristic|llm`) flatten into `extract_mode=profile` or `llm`. Key: `DEEPSEEK_API_KEY` in `.env`. Do not confuse with Cursor dev agents.
## Common tasks
+7 -1
View File
@@ -76,7 +76,8 @@ docker compose up --build
| Раздел | Путь | Описание |
|--------|------|----------|
| **Карта** | `/` | Интерактивная карта событий (публичный просмотр). CRUD ручных объектов — только для админа (ПКМ). Поддерживает `?eventId=` |
| **Парсеры** | `/parsers` | Адаптеры `telegram` / `crawl4ai` / `viina`, интервал, CRUD; дедуп по `source_url` |
| **Парсеры** | `/parsers` | Связка канал + профиль (`heuristic`→`extract_mode=profile`, `llm`→`extract_mode=llm`); дедуп по `source_url` |
| **Профили** | `/parser-profiles` | Reusable heuristic (правила) или llm (instruction/schema); целевые поля = Event |
| **События** | `/events` | Фильтрация, пагинация, просмотр деталей, ссылка «На карте» для событий с координатами |
| **Аналитика** | `/analytics` | KPI-карточки, график динамики ingest за 30 дней, топ населённых пунктов и регионов |
| **ПИ** | `/consumers` | CRUD подписчиков distribution API, ротация ключей, тест среза через `/api/v1/events` |
@@ -214,3 +215,8 @@ docker compose down
```
Данные PostgreSQL сохраняются в volume `pgdata`.
## Вход в админку (/login):
Логин: admin
Пароль: change-me
+3
View File
@@ -98,8 +98,11 @@ class ParserProfile(Base):
id: Mapped[int] = mapped_column(Integer, primary_key=True, index=True)
name: Mapped[str] = mapped_column(String(255), nullable=False)
# heuristic = static rules; llm = instruction + extract_schema at CP runtime
kind: Mapped[str] = mapped_column(String(50), default="heuristic", nullable=False, index=True)
sample_post: Mapped[str] = mapped_column(Text, default="")
heuristic_profile: Mapped[dict | None] = mapped_column(JSON, nullable=True)
llm_profile: Mapped[dict | None] = mapped_column(JSON, nullable=True)
status: Mapped[str] = mapped_column(String(50), default="draft", index=True)
created_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True),
+17 -3
View File
@@ -87,7 +87,11 @@ def create_parse_job(payload: ParseJobCreate, db: Session = Depends(get_db)):
raise HTTPException(status_code=404, detail="Profile not found")
if not channel:
raise HTTPException(status_code=404, detail="Channel not found")
if not profile.heuristic_profile:
kind = (profile.kind or "heuristic").strip().lower()
if kind == "llm":
if not profile.llm_profile:
raise HTTPException(status_code=400, detail="LLM profile has no llm_profile")
elif not profile.heuristic_profile:
raise HTTPException(status_code=400, detail="Profile has no heuristic_profile")
if not channel.is_active:
raise HTTPException(status_code=400, detail="Channel is inactive")
@@ -146,7 +150,10 @@ def retry_parse_job(job_id: int, db: Session = Depends(get_db)):
job = db.query(ParseJob).filter(ParseJob.id == job_id).first()
if not job:
raise HTTPException(status_code=404, detail="Job not found")
if job.status in ("queued", "running"):
from ..services.job_stale import is_stale_job
if job.status in ("queued", "running") and not is_stale_job(job):
raise HTTPException(status_code=409, detail="Job is already running or queued")
job.status = "queued"
@@ -218,7 +225,14 @@ def delete_parse_job(job_id: int, db: Session = Depends(get_db)):
if not job:
raise HTTPException(status_code=404, detail="Job not found")
if job.status == "running":
raise HTTPException(status_code=409, detail="Cannot delete a running job")
from ..services.job_stale import is_stale_job
if not is_stale_job(job):
raise HTTPException(status_code=409, detail="Cannot delete a running job")
# Stale running — allow delete after marking failed for audit trail
job.status = "failed"
job.last_error = "Deleted while stale running"
db.commit()
db.delete(job)
db.commit()
@@ -52,7 +52,8 @@ def update_job_status(
raise HTTPException(status_code=404, detail="Job not found")
job.status = status
if status in ("completed", "failed"):
# Anchor staleness detection: running/queued start, and terminal finish
if status in ("running", "queued", "completed", "failed"):
job.last_run_at = datetime.now(timezone.utc)
job.last_error = error
db.commit()
@@ -9,6 +9,7 @@ from pydantic import BaseModel, Field
from sqlalchemy.orm import Session
from contracts.heuristic_profile import HeuristicProfile
from contracts.llm_profile import LlmProfile
from ..database import get_db
from ..deps import verify_admin
@@ -44,6 +45,22 @@ class PreviewRequest(BaseModel):
class PreviewResponse(BaseModel):
fields: dict[str, str]
matched: bool = True
missing_required: list[str] = Field(default_factory=list)
class PreviewLlmRequest(BaseModel):
sample_post: str = Field(min_length=1)
llm_profile: dict[str, Any]
class PreviewLlmResponse(BaseModel):
fields: dict[str, str]
events: list[dict[str, str]] = Field(default_factory=list)
matched: bool = True
missing_required: list[str] = Field(default_factory=list)
is_event: bool = True
matched_count: int = 0
def _validate_heuristic_profile(raw: dict[str, Any] | None) -> dict[str, Any] | None:
@@ -55,9 +72,33 @@ def _validate_heuristic_profile(raw: dict[str, Any] | None) -> dict[str, Any] |
raise HTTPException(status_code=400, detail=f"Invalid heuristic_profile: {exc}") from exc
def _profile_status(heuristic_profile: dict | None, explicit: str | None = None) -> str:
def _validate_llm_profile(raw: dict[str, Any] | None) -> dict[str, Any] | None:
if raw is None:
return None
try:
return LlmProfile.model_validate(raw).model_dump()
except Exception as exc:
raise HTTPException(status_code=400, detail=f"Invalid llm_profile: {exc}") from exc
def _normalize_kind(kind: str | None) -> str:
value = (kind or "heuristic").strip().lower()
if value not in ("heuristic", "llm"):
raise HTTPException(status_code=400, detail="kind must be heuristic or llm")
return value
def _profile_status(
*,
kind: str,
heuristic_profile: dict | None,
llm_profile: dict | None,
explicit: str | None = None,
) -> str:
if explicit:
return explicit
if kind == "llm":
return "ready" if llm_profile else "draft"
return "ready" if heuristic_profile else "draft"
@@ -66,6 +107,18 @@ def target_fields():
return {"fields": builder.get_target_fields()}
@router.get("/llm-defaults")
def llm_defaults():
from contracts.llm_profile import DEFAULT_EXTRACT_SCHEMA, DEFAULT_INSTRUCTION
return {
"instruction": DEFAULT_INSTRUCTION,
"extract_schema": DEFAULT_EXTRACT_SCHEMA,
"required_fields": [],
"multi_event": False,
}
@router.post("/generate", response_model=GenerateResponse)
async def generate_profile(payload: GenerateRequest):
if not builder.deepseek_enabled():
@@ -87,6 +140,9 @@ async def generate_profile(payload: GenerateRequest):
detail=f"DeepSeek generate failed: {exc}",
) from exc
dumped = profile.model_dump()
if payload.current_profile and isinstance(payload.current_profile.get("required_fields"), list):
dumped["required_fields"] = payload.current_profile["required_fields"]
dumped = HeuristicProfile.model_validate(dumped).model_dump()
preview = builder.preview_with_profile(payload.sample_post, dumped)
return GenerateResponse(
profile=dumped,
@@ -102,10 +158,40 @@ def preview_profile(payload: PreviewRequest):
raise HTTPException(status_code=400, detail="heuristic_profile is required")
try:
HeuristicProfile.model_validate(raw)
fields = builder.preview_with_profile(payload.sample_post, raw)
fields, matched, missing = builder.match_preview(payload.sample_post, raw)
except Exception as exc:
raise HTTPException(status_code=400, detail=str(exc)) from exc
return PreviewResponse(fields=fields)
return PreviewResponse(fields=fields, matched=matched, missing_required=missing)
@router.post("/preview-llm", response_model=PreviewLlmResponse)
async def preview_llm_profile(payload: PreviewLlmRequest):
if not builder.deepseek_enabled():
raise HTTPException(
status_code=503,
detail="DEEPSEEK_API_KEY is not set on ca-api. Add it to .env for LLM preview.",
)
try:
LlmProfile.model_validate(payload.llm_profile)
events, fields, matched, missing, is_event, matched_count = await builder.preview_llm_extract(
payload.sample_post,
payload.llm_profile,
)
except ValueError as exc:
raise HTTPException(status_code=400, detail=str(exc)) from exc
except Exception as exc:
raise HTTPException(
status_code=502,
detail=f"DeepSeek LLM preview failed: {exc}",
) from exc
return PreviewLlmResponse(
fields=fields,
events=events,
matched=matched,
missing_required=missing,
is_event=is_event,
matched_count=matched_count,
)
@router.get("", response_model=list[ParserProfileRead])
@@ -115,12 +201,26 @@ def list_profiles(db: Session = Depends(get_db)):
@router.post("", response_model=ParserProfileRead, status_code=201)
def create_profile(payload: ParserProfileCreate, db: Session = Depends(get_db)):
kind = _normalize_kind(payload.kind)
heuristic = _validate_heuristic_profile(payload.heuristic_profile)
llm = _validate_llm_profile(payload.llm_profile)
if kind == "heuristic" and llm and not heuristic:
# ignore stray llm blob when creating heuristic
llm = None
if kind == "llm" and heuristic and not llm:
heuristic = None
profile = ParserProfile(
name=payload.name.strip(),
kind=kind,
sample_post=payload.sample_post or "",
heuristic_profile=heuristic,
status=_profile_status(heuristic, payload.status),
heuristic_profile=heuristic if kind == "heuristic" else None,
llm_profile=llm if kind == "llm" else None,
status=_profile_status(
kind=kind,
heuristic_profile=heuristic if kind == "heuristic" else None,
llm_profile=llm if kind == "llm" else None,
explicit=payload.status,
),
)
db.add(profile)
db.commit()
@@ -150,12 +250,33 @@ def update_profile(
profile.name = payload.name.strip()
if payload.sample_post is not None:
profile.sample_post = payload.sample_post
if payload.kind is not None:
profile.kind = _normalize_kind(payload.kind)
kind = _normalize_kind(profile.kind)
if "heuristic_profile" in payload.model_fields_set:
profile.heuristic_profile = _validate_heuristic_profile(payload.heuristic_profile)
if "llm_profile" in payload.model_fields_set:
profile.llm_profile = _validate_llm_profile(payload.llm_profile)
# Keep only the blob matching kind
if kind == "heuristic":
profile.llm_profile = None
else:
profile.heuristic_profile = None
if payload.status is not None:
profile.status = payload.status
elif "heuristic_profile" in payload.model_fields_set:
profile.status = _profile_status(profile.heuristic_profile)
elif (
"heuristic_profile" in payload.model_fields_set
or "llm_profile" in payload.model_fields_set
or payload.kind is not None
):
profile.status = _profile_status(
kind=kind,
heuristic_profile=profile.heuristic_profile,
llm_profile=profile.llm_profile,
)
db.commit()
db.refresh(profile)
+6
View File
@@ -161,15 +161,19 @@ class IngestResponse(BaseModel):
class ParserProfileCreate(BaseModel):
name: str = Field(min_length=1, max_length=255)
kind: str = Field(default="heuristic", pattern="^(heuristic|llm)$")
sample_post: str = ""
heuristic_profile: dict[str, Any] | None = None
llm_profile: dict[str, Any] | None = None
status: str | None = None
class ParserProfileUpdate(BaseModel):
name: str | None = Field(default=None, min_length=1, max_length=255)
kind: str | None = Field(default=None, pattern="^(heuristic|llm)$")
sample_post: str | None = None
heuristic_profile: dict[str, Any] | None = None
llm_profile: dict[str, Any] | None = None
status: str | None = None
@@ -178,8 +182,10 @@ class ParserProfileRead(BaseModel):
id: int
name: str
kind: str = "heuristic"
sample_post: str
heuristic_profile: dict[str, Any] | None
llm_profile: dict[str, Any] | None = None
status: str
created_at: datetime
@@ -0,0 +1,36 @@
"""Stale parse-job recovery helpers (running/queued left behind after worker crash)."""
from __future__ import annotations
import os
from datetime import datetime, timezone
from ..models import ParseJob
# Jobs stuck in running/queued longer than this are considered abandoned.
STALE_JOB_SECONDS = int(os.getenv("STALE_JOB_SECONDS", "900"))
def _aware(dt: datetime | None) -> datetime | None:
if dt is None:
return None
if dt.tzinfo is None:
return dt.replace(tzinfo=timezone.utc)
return dt
def job_anchor_time(job: ParseJob) -> datetime | None:
"""Best available timestamp for staleness (prefer last_run_at)."""
return _aware(job.last_run_at) or _aware(getattr(job, "created_at", None))
def is_stale_job(job: ParseJob, now: datetime | None = None, *, ttl: int | None = None) -> bool:
if job.status not in ("running", "queued"):
return False
now = now or datetime.now(timezone.utc)
anchor = job_anchor_time(job)
if anchor is None:
# No timestamp — treat long-lived running as stale immediately for recovery
return job.status == "running"
limit = ttl if ttl is not None else STALE_JOB_SECONDS
return (now - anchor).total_seconds() >= limit
@@ -33,6 +33,25 @@ def flatten_pair_config(
limit: int = 100,
) -> dict:
"""Expand Profile + Channel into Redis/CP source_config."""
kind = (profile.kind or "heuristic").strip().lower()
if kind == "llm":
from contracts.llm_profile import LlmProfile
if not profile.llm_profile:
raise ValueError("ParserProfile.llm_profile is empty")
llm = LlmProfile.model_validate(profile.llm_profile)
cfg = TelegramSourceConfig(
channel=channel.channel.strip(),
limit=limit,
extract_mode="llm",
extract_schema=llm.extract_schema,
instruction=llm.instruction,
required_fields=list(llm.required_fields),
multi_event=bool(llm.multi_event),
sample_post=profile.sample_post or None,
)
return cfg.model_dump()
if not profile.heuristic_profile:
raise ValueError("ParserProfile.heuristic_profile is empty")
cfg = TelegramSourceConfig(
@@ -23,11 +23,14 @@ def migrate_schema(engine: Engine) -> None:
if "channel_id" not in columns:
statements.append("ALTER TABLE parse_jobs ADD COLUMN channel_id INTEGER")
# create_all handles new tables; FKs on existing DBs may need indexes
if "parse_jobs" in tables:
# Re-inspect after potential adds is not needed for FK constraints here —
# create_all + nullable FKs are enough for MVP; optional constraints below.
pass
if "parser_profiles" in tables:
columns = {col["name"] for col in inspector.get_columns("parser_profiles")}
if "kind" not in columns:
statements.append(
"ALTER TABLE parser_profiles ADD COLUMN kind VARCHAR(50) NOT NULL DEFAULT 'heuristic'"
)
if "llm_profile" not in columns:
statements.append("ALTER TABLE parser_profiles ADD COLUMN llm_profile JSON")
if not statements:
return
@@ -14,6 +14,7 @@ from contracts.heuristic_profile import (
HeuristicProfile,
TARGET_FIELDS,
apply_profile,
match_profile,
target_field_specs,
)
@@ -47,6 +48,14 @@ def preview_with_profile(sample_post: str, profile: dict[str, Any] | HeuristicPr
return apply_profile(sample_post, profile)
def match_preview(
sample_post: str,
profile: dict[str, Any] | HeuristicProfile,
) -> tuple[dict[str, str], bool, list[str]]:
matched, fields, missing = match_profile(sample_post, profile)
return fields, matched, missing
def empty_preview_fields(preview: dict[str, str]) -> list[str]:
return [name for name, value in preview.items() if not (value or "").strip()]
@@ -191,3 +200,111 @@ async def generate_profile(
content = data["choices"][0]["message"]["content"]
return _parse_profile_response(content)
async def preview_llm_extract(
sample_post: str,
llm_profile: dict[str, Any],
) -> tuple[list[dict[str, str]], dict[str, str], bool, list[str], bool, int]:
"""One-shot DeepSeek extract for CA preview (does not persist).
Returns (events, fields, matched, missing_required, is_event, matched_count).
``fields`` is the first event (or empty) for legacy UI compatibility.
"""
from contracts.llm_profile import DEFAULT_INSTRUCTION, LlmProfile, match_llm_required
settings = deepseek_settings()
if not settings["api_key"]:
raise RuntimeError(
"DEEPSEEK_API_KEY is not set on ca-api. Add it to .env for LLM preview."
)
sample = (sample_post or "").strip()
if not sample:
raise ValueError("sample_post is required")
profile = LlmProfile.model_validate(llm_profile)
schema = profile.extract_schema
instr = profile.instruction or DEFAULT_INSTRUCTION
schema_lines = "\n".join(f"- {k}: {v}" for k, v in schema.items())
multi = bool(profile.multi_event)
if multi:
user_prompt = (
f"{instr}\n\n"
"If the text describes multiple distinct events (different places, "
"coords, or dates), return one object per event in \"events\".\n"
f"Fields per event:\n{schema_lines}\n\n"
'Return JSON: {"is_event": true|false, "events": [{<field>: <string>}, ...]}\n'
"If there is no event, return is_event=false and events=[].\n\n"
f"Text:\n{sample[:12000]}"
)
else:
user_prompt = (
f"{instr}\n\n"
f"Fields to extract:\n{schema_lines}\n\n"
'Return JSON: {"is_event": true|false, "fields": {<field>: <string>}}\n\n'
f"Text:\n{sample[:12000]}"
)
payload = {
"model": settings["model"],
"messages": [
{
"role": "system",
"content": (
"You extract structured event data for a geoint map. "
"Output valid JSON only, no markdown."
),
},
{"role": "user", "content": user_prompt},
],
"temperature": 0.1,
"response_format": {"type": "json_object"},
}
url = f"{settings['base_url']}/chat/completions"
async with httpx.AsyncClient(timeout=90.0) as client:
response = await client.post(
url,
headers={
"Authorization": f"Bearer {settings['api_key']}",
"Content-Type": "application/json",
},
json=payload,
)
response.raise_for_status()
data = response.json()
content = data["choices"][0]["message"]["content"]
parsed = _extract_json_object(content)
is_event = bool(parsed.get("is_event", True))
empty_fields = {key: "" for key in schema}
if not is_event:
return [], empty_fields, False, [], False, 0
events: list[dict[str, str]] = []
if multi and isinstance(parsed.get("events"), list):
for item in parsed["events"]:
if isinstance(item, dict):
events.append({key: str(item.get(key) or "").strip() for key in schema})
else:
fields_raw = parsed.get("fields") if isinstance(parsed.get("fields"), dict) else parsed
if not isinstance(fields_raw, dict):
fields_raw = {}
events.append({key: str(fields_raw.get(key) or "").strip() for key in schema})
if not events:
return [], empty_fields, False, [], False, 0
matched_events = [ev for ev in events if match_llm_required(ev, list(profile.required_fields))]
fields = events[0]
missing = [
name
for name in profile.required_fields
if not str(fields.get(name) or "").strip()
]
matched = match_llm_required(fields, list(profile.required_fields))
return events, fields, matched, missing, True, len(matched_events)
@@ -5,6 +5,7 @@ from datetime import datetime, timezone
from ..database import SessionLocal
from ..models import ParseJob
from .job_stale import STALE_JOB_SECONDS, is_stale_job
from .jobs import enqueue_parse_job
logger = logging.getLogger(__name__)
@@ -13,10 +14,47 @@ TICK_SECONDS = 30
RECURRING_STATUSES = ("completed", "failed")
def recover_stale_jobs(db, now: datetime) -> None:
"""Mark abandoned running/queued jobs as failed; re-queue if still active."""
stuck = (
db.query(ParseJob)
.filter(ParseJob.status.in_(("running", "queued")))
.all()
)
for job in stuck:
if not is_stale_job(job, now):
continue
prev = job.status
job.status = "failed"
job.last_error = (
f"Stale {prev} recovered after {STALE_JOB_SECONDS}s "
"(worker likely restarted)"
)
job.last_run_at = now
db.commit()
logger.warning("Recovered stale job %s (was %s)", job.id, prev)
if not job.is_active or job.interval_seconds <= 0:
continue
job.status = "queued"
job.last_error = None
db.commit()
try:
enqueue_parse_job(db, job)
logger.info("Re-queued recovered job %s", job.id)
except ValueError as exc:
job.status = "failed"
job.last_error = str(exc)
db.commit()
logger.warning("Skip re-queue recovered job %s: %s", job.id, exc)
def run_scheduler_tick() -> None:
db = SessionLocal()
try:
now = datetime.now(timezone.utc)
recover_stale_jobs(db, now)
jobs = (
db.query(ParseJob)
.filter(
@@ -61,5 +99,9 @@ def start_scheduler() -> threading.Event:
stop_event = threading.Event()
thread = threading.Thread(target=_scheduler_loop, args=(stop_event,), daemon=True)
thread.start()
logger.info("Parse job scheduler started (tick every %ss)", TICK_SECONDS)
logger.info(
"Parse job scheduler started (tick every %ss, stale after %ss)",
TICK_SECONDS,
STALE_JOB_SECONDS,
)
return stop_event
+33 -3
View File
@@ -1,6 +1,7 @@
import { ADMIN_BASE, request } from "./client";
import type {
HeuristicProfile,
LlmProfile,
ParseChannel,
ParseChannelCreate,
ParseChannelUpdate,
@@ -10,7 +11,7 @@ import type {
TargetField,
} from "../types/admin";
export type { HeuristicProfile, TargetField } from "../types/admin";
export type { HeuristicProfile, LlmProfile, TargetField } from "../types/admin";
export function fetchChannels(): Promise<ParseChannel[]> {
return request<ParseChannel[]>("/parse-channels", undefined, ADMIN_BASE);
@@ -64,6 +65,10 @@ export function fetchTargetFields(): Promise<{ fields: TargetField[] }> {
return request<{ fields: TargetField[] }>("/parser-profiles/target-fields", undefined, ADMIN_BASE);
}
export function fetchLlmDefaults(): Promise<LlmProfile> {
return request<LlmProfile>("/parser-profiles/llm-defaults", undefined, ADMIN_BASE);
}
export function generateParserProfile(payload: {
sample_post: string;
hint?: string;
@@ -94,10 +99,35 @@ export function generateParserProfile(payload: {
export function previewParserProfile(
sample_post: string,
heuristic_profile: HeuristicProfile | Record<string, unknown>,
): Promise<{ fields: Record<string, string> }> {
return request<{ fields: Record<string, string> }>(
): Promise<{ fields: Record<string, string>; matched: boolean; missing_required: string[] }> {
return request<{ fields: Record<string, string>; matched: boolean; missing_required: string[] }>(
"/parser-profiles/preview",
{ method: "POST", body: JSON.stringify({ sample_post, heuristic_profile }) },
ADMIN_BASE,
);
}
export function previewLlmProfile(
sample_post: string,
llm_profile: LlmProfile | Record<string, unknown>,
): Promise<{
fields: Record<string, string>;
events: Record<string, string>[];
matched: boolean;
missing_required: string[];
is_event: boolean;
matched_count: number;
}> {
return request<{
fields: Record<string, string>;
events: Record<string, string>[];
matched: boolean;
missing_required: string[];
is_event: boolean;
matched_count: number;
}>(
"/parser-profiles/preview-llm",
{ method: "POST", body: JSON.stringify({ sample_post, llm_profile }) },
ADMIN_BASE,
);
}
+26 -12
View File
@@ -36,26 +36,52 @@ export interface ParseJobUpdate {
is_active?: boolean;
}
export type TargetField = {
name: string;
type: string;
description: string;
};
export type HeuristicProfile = {
version: number;
fields: Record<string, Record<string, unknown>>;
required_fields?: string[];
notes?: string;
};
export type LlmProfile = {
instruction?: string | null;
extract_schema: Record<string, string>;
required_fields?: string[];
multi_event?: boolean;
};
export interface ParserProfile {
id: number;
name: string;
kind: "heuristic" | "llm";
sample_post: string;
heuristic_profile: Record<string, unknown> | null;
llm_profile: LlmProfile | Record<string, unknown> | null;
status: string;
created_at: string;
}
export interface ParserProfileCreate {
name: string;
kind?: "heuristic" | "llm";
sample_post?: string;
heuristic_profile?: Record<string, unknown> | null;
llm_profile?: LlmProfile | Record<string, unknown> | null;
status?: string;
}
export interface ParserProfileUpdate {
name?: string;
kind?: "heuristic" | "llm";
sample_post?: string;
heuristic_profile?: Record<string, unknown> | null;
llm_profile?: LlmProfile | Record<string, unknown> | null;
status?: string;
}
@@ -82,18 +108,6 @@ export interface ParseChannelUpdate {
is_active?: boolean;
}
export type TargetField = {
name: string;
type: string;
description: string;
};
export type HeuristicProfile = {
version: number;
fields: Record<string, Record<string, unknown>>;
notes?: string;
};
export interface EventRecord {
id: number;
source_type: string;
@@ -1,14 +1,17 @@
<script setup lang="ts">
import { onMounted, ref } from "vue";
import { onMounted, ref, watch } from "vue";
import {
createProfile,
deleteProfile,
fetchLlmDefaults,
fetchProfiles,
fetchTargetFields,
generateParserProfile,
previewLlmProfile,
previewParserProfile,
updateProfile,
type HeuristicProfile,
type LlmProfile,
type TargetField,
} from "../api/entities";
import type { ParserProfile } from "../types/admin";
@@ -19,18 +22,29 @@ const error = ref("");
const success = ref("");
const submitting = ref(false);
const profileKind = ref<"heuristic" | "llm">("heuristic");
const samplePost = ref("");
const generateHint = ref("");
const profileName = ref("");
const targetFields = ref<TargetField[]>([]);
const profileJson = ref("");
const llmInstruction = ref("");
const llmSchemaJson = ref("");
const llmMultiEvent = ref(false);
const previewFields = ref<Record<string, string> | null>(null);
const previewEvents = ref<Record<string, string>[]>([]);
const emptyFields = ref<string[]>([]);
const requiredFields = ref<string[]>([]);
const previewMatched = ref(true);
const previewMissing = ref<string[]>([]);
const previewIsEvent = ref(true);
const previewMatchedCount = ref(0);
const generating = ref(false);
const previewing = ref(false);
const loadingFields = ref(true);
const editingId = ref<number | null>(null);
const profileStatus = ref<"draft" | "ready">("draft");
const llmDefaults = ref<LlmProfile | null>(null);
async function loadProfiles() {
loading.value = true;
@@ -56,7 +70,28 @@ async function loadTargetFields() {
}
}
function syncProfileFromJson(): HeuristicProfile | null {
async function loadLlmDefaults() {
try {
llmDefaults.value = await fetchLlmDefaults();
} catch {
llmDefaults.value = null;
}
}
function applyLlmDefaults() {
const d = llmDefaults.value;
if (!d) return;
llmInstruction.value = d.instruction || "";
llmSchemaJson.value = JSON.stringify(d.extract_schema || {}, null, 2);
if (!requiredFields.value.length) {
requiredFields.value = Array.isArray(d.required_fields) ? [...d.required_fields] : [];
}
if (typeof d.multi_event === "boolean") {
llmMultiEvent.value = d.multi_event;
}
}
function syncHeuristicFromJson(): HeuristicProfile | null {
if (!profileJson.value.trim()) return null;
try {
return JSON.parse(profileJson.value) as HeuristicProfile;
@@ -66,7 +101,88 @@ function syncProfileFromJson(): HeuristicProfile | null {
}
}
function syncLlmFromForm(): LlmProfile | null {
let schema: Record<string, string>;
try {
schema = JSON.parse(llmSchemaJson.value || "{}") as Record<string, string>;
} catch {
error.value = "Некорректный JSON extract_schema";
return null;
}
if (!schema || typeof schema !== "object" || !Object.keys(schema).length) {
error.value = "extract_schema не должен быть пустым";
return null;
}
return {
instruction: llmInstruction.value.trim() || null,
extract_schema: schema,
required_fields: [...requiredFields.value],
multi_event: llmMultiEvent.value,
};
}
const DEFAULT_REQUIRED = ["event_date", "coords"];
function uniqueFields(names: string[]): string[] {
return names.filter((name, idx) => names.indexOf(name) === idx);
}
function writeRequiredToHeuristicJson(names: string[]) {
const current = syncHeuristicFromJson();
if (!current) return;
current.required_fields = names;
profileJson.value = JSON.stringify(current, null, 2);
}
function requiredFromHeuristic(
profile: HeuristicProfile,
preview: Record<string, string> | null,
): string[] {
if (Array.isArray(profile.required_fields)) {
return uniqueFields(profile.required_fields);
}
if (!preview) return [];
return DEFAULT_REQUIRED.filter((name) => !!(preview[name] || "").trim());
}
function isRequired(name: string): boolean {
return requiredFields.value.includes(name);
}
async function toggleRequired(name: string) {
const next = isRequired(name)
? requiredFields.value.filter((n) => n !== name)
: [...requiredFields.value, name];
requiredFields.value = next;
if (profileKind.value === "heuristic") {
writeRequiredToHeuristicJson(next);
}
if (previewFields.value && samplePost.value.trim()) {
await handlePreview();
}
}
watch(profileKind, (kind, prev) => {
if (kind === prev) return;
previewFields.value = null;
previewEvents.value = [];
emptyFields.value = [];
previewMatched.value = true;
previewMissing.value = [];
previewIsEvent.value = true;
previewMatchedCount.value = 0;
success.value = "";
if (kind === "llm" && !llmSchemaJson.value.trim()) {
applyLlmDefaults();
}
});
async function handleGenerate() {
if (profileKind.value === "llm") {
applyLlmDefaults();
success.value = "Подставлены дефолты LLM schema/instruction. Отредактируйте и Preview.";
return;
}
if (!samplePost.value.trim()) {
error.value = "Вставьте образец поста";
return;
@@ -79,7 +195,7 @@ async function handleGenerate() {
try {
let current: HeuristicProfile | null = null;
if (profileJson.value.trim()) {
current = syncProfileFromJson();
current = syncHeuristicFromJson();
if (!current) return;
}
const data = await generateParserProfile({
@@ -87,9 +203,18 @@ async function handleGenerate() {
hint: generateHint.value.trim() || undefined,
current_profile: current,
});
const required =
current && Array.isArray(current.required_fields)
? uniqueFields(current.required_fields)
: DEFAULT_REQUIRED.filter((name) => !!(data.preview[name] || "").trim());
data.profile.required_fields = required;
profileJson.value = JSON.stringify(data.profile, null, 2);
requiredFields.value = required;
previewFields.value = data.preview;
emptyFields.value = data.empty_fields || [];
const previewData = await previewParserProfile(samplePost.value, data.profile);
previewMatched.value = previewData.matched;
previewMissing.value = previewData.missing_required || [];
if (emptyFields.value.length) {
success.value =
`Профиль обновлён, но пустые поля: ${emptyFields.value.join(", ")}. ` +
@@ -109,11 +234,6 @@ async function handleGenerate() {
}
async function handlePreview() {
const current = syncProfileFromJson();
if (!current) {
if (!error.value) error.value = "Сначала сгенерируйте или вставьте профиль";
return;
}
if (!samplePost.value.trim()) {
error.value = "Нужен образец поста для preview";
return;
@@ -121,11 +241,53 @@ async function handlePreview() {
previewing.value = true;
error.value = "";
try {
const data = await previewParserProfile(samplePost.value, current);
previewFields.value = data.fields;
emptyFields.value = Object.entries(data.fields)
.filter(([, v]) => !v)
.map(([k]) => k);
if (profileKind.value === "llm") {
const llm = syncLlmFromForm();
if (!llm) return;
const data = await previewLlmProfile(samplePost.value, llm);
const rows =
Array.isArray(data.events) && data.events.length
? data.events
: data.fields
? [data.fields]
: [];
previewEvents.value = rows;
previewFields.value = rows[0] || data.fields || null;
emptyFields.value = previewFields.value
? Object.entries(previewFields.value)
.filter(([, v]) => !v)
.map(([k]) => k)
: [];
previewMatched.value = data.matched;
previewMissing.value = data.missing_required || [];
previewIsEvent.value = data.is_event;
previewMatchedCount.value = data.matched_count ?? rows.length;
if (!data.is_event) {
success.value = "Модель пометила текст как не-событие (is_event=false) — runtime пропустит.";
} else if (rows.length > 1) {
success.value = `Preview: ${rows.length} событий из поста (пройдут фильтр: ${previewMatchedCount.value}).`;
}
} else {
const current = syncHeuristicFromJson();
if (!current) {
if (!error.value) error.value = "Сначала сгенерируйте или вставьте профиль";
return;
}
const data = await previewParserProfile(samplePost.value, current);
previewFields.value = data.fields;
emptyFields.value = Object.entries(data.fields)
.filter(([, v]) => !v)
.map(([k]) => k);
requiredFields.value = requiredFromHeuristic(current, data.fields);
if (!Array.isArray(current.required_fields)) {
writeRequiredToHeuristicJson(requiredFields.value);
}
previewMatched.value = data.matched;
previewMissing.value = data.missing_required || [];
previewIsEvent.value = true;
previewEvents.value = data.fields ? [data.fields] : [];
previewMatchedCount.value = data.matched ? 1 : 0;
}
} catch (err) {
error.value = err instanceof Error ? err.message : "Ошибка preview";
} finally {
@@ -135,29 +297,72 @@ async function handlePreview() {
function resetEditor() {
editingId.value = null;
profileKind.value = "heuristic";
profileName.value = "";
samplePost.value = "";
generateHint.value = "";
profileJson.value = "";
llmInstruction.value = "";
llmSchemaJson.value = "";
llmMultiEvent.value = false;
profileStatus.value = "draft";
previewFields.value = null;
previewEvents.value = [];
emptyFields.value = [];
requiredFields.value = [];
previewMatched.value = true;
previewMissing.value = [];
previewIsEvent.value = true;
previewMatchedCount.value = 0;
success.value = "";
}
function openEdit(p: ParserProfile) {
editingId.value = p.id;
profileKind.value = p.kind === "llm" ? "llm" : "heuristic";
profileName.value = p.name;
samplePost.value = p.sample_post || "";
generateHint.value = "";
profileJson.value = p.heuristic_profile
? JSON.stringify(p.heuristic_profile, null, 2)
: "";
profileStatus.value = p.status === "ready" ? "ready" : "draft";
previewFields.value = null;
previewEvents.value = [];
emptyFields.value = [];
previewMatched.value = true;
previewMissing.value = [];
previewIsEvent.value = true;
previewMatchedCount.value = 0;
success.value = "";
error.value = "";
if (p.kind === "llm") {
profileJson.value = "";
const lp = (p.llm_profile || {}) as LlmProfile;
llmInstruction.value = lp.instruction || "";
llmSchemaJson.value = JSON.stringify(lp.extract_schema || {}, null, 2);
llmMultiEvent.value = !!lp.multi_event;
requiredFields.value = Array.isArray(lp.required_fields)
? uniqueFields(lp.required_fields)
: [];
if (!llmSchemaJson.value.trim() || llmSchemaJson.value === "{}") {
applyLlmDefaults();
}
} else {
llmInstruction.value = "";
llmSchemaJson.value = "";
llmMultiEvent.value = false;
profileJson.value = p.heuristic_profile
? JSON.stringify(p.heuristic_profile, null, 2)
: "";
const hp = p.heuristic_profile;
if (hp && Array.isArray(hp.required_fields)) {
requiredFields.value = uniqueFields(hp.required_fields as string[]);
} else if (hp) {
requiredFields.value = [...DEFAULT_REQUIRED];
writeRequiredToHeuristicJson(requiredFields.value);
} else {
requiredFields.value = [];
}
}
window.scrollTo({ top: 0, behavior: "smooth" });
}
@@ -166,33 +371,60 @@ async function handleSave() {
error.value = "Укажите название профиля";
return;
}
const current = syncProfileFromJson();
if (!current) {
if (!error.value) error.value = "Нужен JSON профиля (Generate или вручную)";
return;
}
submitting.value = true;
error.value = "";
success.value = "";
try {
if (editingId.value != null) {
await updateProfile(editingId.value, {
if (profileKind.value === "llm") {
const llm = syncLlmFromForm();
if (!llm) return;
const payload = {
name: profileName.value.trim(),
kind: "llm" as const,
sample_post: samplePost.value,
heuristic_profile: current,
llm_profile: llm,
heuristic_profile: null,
status: profileStatus.value,
});
success.value = `Профиль #${editingId.value} обновлён`;
};
if (editingId.value != null) {
await updateProfile(editingId.value, payload);
success.value = `LLM-профиль #${editingId.value} обновлён`;
} else {
const created = await createProfile({
...payload,
status: profileStatus.value === "draft" ? "draft" : "ready",
});
success.value = `LLM-профиль #${created.id} сохранён`;
editingId.value = created.id;
profileStatus.value = created.status === "ready" ? "ready" : "draft";
}
} else {
const created = await createProfile({
writeRequiredToHeuristicJson(requiredFields.value);
const current = syncHeuristicFromJson();
if (!current) {
if (!error.value) error.value = "Нужен JSON профиля (Generate или вручную)";
return;
}
const payload = {
name: profileName.value.trim(),
kind: "heuristic" as const,
sample_post: samplePost.value,
heuristic_profile: current,
status: profileStatus.value === "draft" ? "draft" : "ready",
});
success.value = `Профиль #${created.id} сохранён`;
editingId.value = created.id;
profileStatus.value = created.status === "ready" ? "ready" : "draft";
llm_profile: null,
status: profileStatus.value,
};
if (editingId.value != null) {
await updateProfile(editingId.value, payload);
success.value = `Профиль #${editingId.value} обновлён`;
} else {
const created = await createProfile({
...payload,
status: profileStatus.value === "draft" ? "draft" : "ready",
});
success.value = `Профиль #${created.id} сохранён`;
editingId.value = created.id;
profileStatus.value = created.status === "ready" ? "ready" : "draft";
}
}
await loadProfiles();
} catch (err) {
@@ -218,9 +450,18 @@ function formatDate(value: string): string {
return new Date(value).toLocaleString("ru-RU");
}
onMounted(() => {
loadProfiles();
loadTargetFields();
function kindLabel(kind: string | undefined): string {
return kind === "llm" ? "llm" : "heuristic";
}
function canSave(): boolean {
if (!profileName.value.trim()) return false;
if (profileKind.value === "llm") return !!llmSchemaJson.value.trim();
return !!profileJson.value.trim();
}
onMounted(async () => {
await Promise.all([loadProfiles(), loadTargetFields(), loadLlmDefaults()]);
});
</script>
@@ -228,9 +469,10 @@ onMounted(() => {
<div class="page">
<h2 class="page-heading">Профили парсера</h2>
<p class="intro">
Generate строит статичные правила по образцу поста. Если Preview пустой по полю —
напишите подсказку и снова Generate (агент правит текущий JSON). Можно править JSON
вручную. Дальше профиль подключается к каналу на
Два вида профилей под фиксированные поля Event:
<strong>heuristic</strong> — статичные правила (Generate один раз, runtime без LLM);
<strong>llm</strong> — instruction + schema, DeepSeek на каждый пост в batch.
Связка с каналом — на
<router-link to="/parsers">Парсеры</router-link>.
</p>
@@ -246,6 +488,13 @@ onMounted(() => {
Название
<input v-model="profileName" type="text" placeholder="Сводка: дата + НП + coords" />
</label>
<label>
Тип
<select v-model="profileKind" :disabled="editingId != null">
<option value="heuristic">heuristic (правила)</option>
<option value="llm">llm (DeepSeek runtime)</option>
</select>
</label>
<label>
Статус
<select v-model="profileStatus">
@@ -263,35 +512,64 @@ onMounted(() => {
placeholder="15.03.2024 Населённый пункт&#10;&#10;Описание…&#10;&#10;48.123456, 37.654321"
/>
</label>
<label>
Подсказка агенту
<span class="muted"> (если Generate ошибся — опишите, что исправить, и нажмите Generate снова)</span>
<textarea
v-model="generateHint"
class="hint-area"
rows="3"
placeholder="например: description — абзацы между заголовком и координатами, без #хештегов"
/>
</label>
<template v-if="profileKind === 'heuristic'">
<label>
Подсказка агенту
<span class="muted"> (если Generate ошибся — опишите, что исправить, и нажмите Generate снова)</span>
<textarea
v-model="generateHint"
class="hint-area"
rows="3"
placeholder="например: description — абзацы между заголовком и координатами, без #хештегов"
/>
</label>
</template>
<template v-else>
<label>
Instruction
<textarea
v-model="llmInstruction"
class="hint-area"
rows="3"
placeholder="Инструкция для DeepSeek на каждый пост"
/>
</label>
<label class="required-item multi-flag">
<input v-model="llmMultiEvent" type="checkbox" />
Несколько событий в посте
<span class="muted"> (LLM вернёт массив; ingest с URL #e1, #e2…)</span>
</label>
</template>
<div class="form-actions">
<button
type="button"
class="btn btn-primary"
:disabled="generating || !samplePost.trim()"
:disabled="generating || (profileKind === 'heuristic' && !samplePost.trim())"
@click="handleGenerate"
>
{{
generating
? "Генерация..."
: generateHint.trim() || profileJson.trim()
? "Generate / Refine"
: "Generate"
}}
<template v-if="profileKind === 'llm'">
{{ generating ? "…" : "Подставить дефолты schema" }}
</template>
<template v-else>
{{
generating
? "Генерация..."
: generateHint.trim() || profileJson.trim()
? "Generate / Refine"
: "Generate"
}}
</template>
</button>
<button
type="button"
class="btn"
:disabled="previewing || !profileJson.trim()"
:disabled="
previewing ||
(profileKind === 'heuristic' ? !profileJson.trim() : !llmSchemaJson.trim())
"
@click="handlePreview"
>
{{ previewing ? "Preview..." : "Preview" }}
@@ -299,7 +577,7 @@ onMounted(() => {
<button
type="button"
class="btn btn-primary"
:disabled="submitting || !profileJson.trim() || !profileName.trim()"
:disabled="submitting || !canSave()"
@click="handleSave"
>
{{ submitting ? "Сохранение..." : editingId != null ? "Обновить" : "Сохранить профиль" }}
@@ -332,43 +610,120 @@ onMounted(() => {
</section>
<section class="card">
<h3>Профиль (JSON)</h3>
<h3 v-if="profileKind === 'heuristic'">Профиль (JSON)</h3>
<h3 v-else>extract_schema (JSON)</h3>
<textarea
v-if="profileKind === 'heuristic'"
v-model="profileJson"
class="profile-area"
rows="16"
spellcheck="false"
placeholder="Появится после Generate"
/>
<textarea
v-else
v-model="llmSchemaJson"
class="profile-area"
rows="16"
spellcheck="false"
placeholder="Ключ → описание поля для LLM"
/>
</section>
</div>
<section v-if="previewFields" class="card">
<h3>Preview строки</h3>
<h3>
Preview
<span v-if="previewEvents.length > 1" class="muted">
({{ previewEvents.length }} событий)
</span>
<span v-else>строки</span>
</h3>
<p v-if="profileKind === 'llm' && !previewIsEvent" class="form-warn">
is_event=false — пост не будет ингеститься.
</p>
<p v-if="profileKind === 'llm' && previewEvents.length > 1" class="form-success match-ok">
Из поста извлечено {{ previewEvents.length }} событий;
пройдут required_fields: {{ previewMatchedCount }}.
Runtime URL: пост#e1 … #e{{ previewEvents.length }}.
</p>
<p v-if="emptyFields.length" class="form-warn">
Пустые поля: <code>{{ emptyFields.join(", ") }}</code> — уточните подсказку и Generate / Refine.
Пустые поля (первая строка): <code>{{ emptyFields.join(", ") }}</code>
<template v-if="profileKind === 'heuristic'">
— уточните подсказку и Generate / Refine.
</template>
</p>
<p v-if="!requiredFields.length" class="form-warn">
Нет обязательных полей —
<template v-if="profileKind === 'llm'">
фильтр только по is_event.
</template>
<template v-else>
любой пост канала будет сохранён. Отметьте поля, без которых пост нужно отбрасывать.
</template>
</p>
<p v-else-if="!previewMatched && previewEvents.length <= 1" class="form-warn">
Этот образец не прошёл бы фильтр. Не хватает:
<code>{{ previewMissing.join(", ") }}</code>
</p>
<p v-else-if="previewMatched" class="form-success match-ok">
<template v-if="previewEvents.length <= 1">
Образец проходит фильтр. Runtime сохранит только посты с заполненными
<code>{{ requiredFields.join(", ") }}</code>.
</template>
<template v-else>
Обязательные поля: <code>{{ requiredFields.join(", ") }}</code>
(проверка на каждое событие).
</template>
</p>
<div class="table-wrap">
<table class="admin-table">
<thead>
<tr>
<th v-if="previewEvents.length > 1">#</th>
<th v-for="key in Object.keys(previewFields)" :key="key">{{ key }}</th>
</tr>
</thead>
<tbody>
<tr>
<tr v-for="(row, idx) in (previewEvents.length ? previewEvents : [previewFields])" :key="idx">
<td v-if="previewEvents.length > 1">{{ idx + 1 }}</td>
<td
v-for="(val, key) in previewFields"
v-for="key in Object.keys(previewFields)"
:key="key"
class="preview-cell"
:class="{ empty: !val }"
:class="{ empty: !row?.[key] }"
>
{{ val || "—" }}
{{ row?.[key] || "—" }}
</td>
</tr>
</tbody>
</table>
</div>
<div class="required-box">
<h4>Обязательно для сохранения</h4>
<p class="muted required-hint">
<template v-if="profileKind === 'llm'">
После LLM: событие отбрасывается, если обязательное поле пустое (поверх is_event).
</template>
<template v-else>
Пост отбрасывается, если хоть одно отмеченное поле не извлеклось (для coords — ещё и не парсится lat,lon).
</template>
</p>
<div class="required-grid">
<label
v-for="key in Object.keys(previewFields)"
:key="key"
class="required-item"
>
<input
type="checkbox"
:checked="isRequired(key)"
@change="toggleRequired(key)"
/>
<code>{{ key }}</code>
</label>
</div>
</div>
</section>
<p v-if="error" class="form-error">{{ error }}</p>
@@ -382,6 +737,7 @@ onMounted(() => {
<tr>
<th>ID</th>
<th>Название</th>
<th>Тип</th>
<th>Статус</th>
<th>Создан</th>
<th></th>
@@ -391,6 +747,7 @@ onMounted(() => {
<tr v-for="p in profiles" :key="p.id">
<td>{{ p.id }}</td>
<td>{{ p.name }}</td>
<td><code>{{ kindLabel(p.kind) }}</code></td>
<td>{{ p.status }}</td>
<td>{{ formatDate(p.created_at) }}</td>
<td class="actions">
@@ -401,7 +758,7 @@ onMounted(() => {
</td>
</tr>
<tr v-if="!loading && !profiles.length">
<td colspan="5" class="muted">Пока нет профилей</td>
<td colspan="6" class="muted">Пока нет профилей</td>
</tr>
</tbody>
</table>
@@ -486,6 +843,42 @@ onMounted(() => {
margin: 0 0 0.75rem;
}
.match-ok {
margin: 0 0 0.75rem;
}
.required-box {
margin-top: 1rem;
}
.required-box h4 {
margin: 0 0 0.25rem;
font-size: 0.9rem;
}
.required-hint {
margin: 0 0 0.6rem;
font-size: 0.8rem;
}
.required-grid {
display: flex;
flex-wrap: wrap;
gap: 0.5rem 1rem;
}
.required-item {
display: flex;
align-items: center;
gap: 0.35rem;
font-size: 0.85rem;
}
.multi-flag {
margin: 0.75rem 0 0;
flex-wrap: wrap;
}
.actions {
display: flex;
gap: 0.35rem;
@@ -38,7 +38,11 @@ const editForm = ref({
});
const readyProfiles = computed(() =>
profiles.value.filter((p) => !!p.heuristic_profile),
profiles.value.filter((p) => {
if (p.status !== "ready") return false;
if (p.kind === "llm") return !!p.llm_profile;
return !!p.heuristic_profile;
}),
);
const activeChannels = computed(() => channels.value.filter((c) => c.is_active));
@@ -186,8 +190,17 @@ async function handleDelete(job: ParseJob) {
}
}
const STALE_JOB_MS = 15 * 60 * 1000;
function isStaleJob(job: ParseJob): boolean {
if (job.status !== "queued" && job.status !== "running") return false;
if (!job.last_run_at) return job.status === "running";
return Date.now() - new Date(job.last_run_at).getTime() >= STALE_JOB_MS;
}
function canForceRun(job: ParseJob): boolean {
return job.status !== "queued" && job.status !== "running";
if (job.status !== "queued" && job.status !== "running") return true;
return isStaleJob(job);
}
async function handleForceRun(jobId: number) {
@@ -237,9 +250,10 @@ onUnmounted(() => {
Создайте
<router-link to="/channels">канал</router-link>
и
<router-link to="/parser-profiles">профиль</router-link>,
затем свяжите их здесь. В Redis уходит плоский
<code>extract_mode=profile</code> + <code>heuristic_profile</code>.
<router-link to="/parser-profiles">профиль</router-link>
(heuristic или llm), затем свяжите их здесь. Flatten в Redis:
heuristic → <code>extract_mode=profile</code>;
llm → <code>extract_mode=llm</code> (batch; listener для llm пока fallback).
</p>
<form class="admin-form" @submit.prevent="handleCreatePair">
<div class="form-row">
@@ -257,7 +271,7 @@ onUnmounted(() => {
<select v-model.number="pairForm.profile_id" required>
<option :value="null" disabled>Выберите…</option>
<option v-for="p in readyProfiles" :key="p.id" :value="p.id">
{{ p.name }} (#{{ p.id }})
{{ p.name }} (#{{ p.id }}, {{ p.kind === "llm" ? "llm" : "heuristic" }})
</option>
</select>
</label>
@@ -343,7 +357,7 @@ onUnmounted(() => {
<button
class="btn btn-sm btn-danger"
type="button"
:disabled="job.status === 'running'"
:disabled="job.status === 'running' && !isStaleJob(job)"
@click="handleDelete(job)"
>
Удалить
+28 -6
View File
@@ -54,14 +54,29 @@ async def run(job_id, source_config, *, ctx) -> tuple[list[dict], str | None]
- `source_type` события = тип адаптера;
- `source_config` валидируется схемами из `contracts/sources.py`.
### Режимы извлечения (Telegram)
Цель всегда одна: фиксированные поля `IngestEventItem` / Event в ЦА (`title`, `description`, `locality`, `coords`→lat/lng, `event_date`, `topic`, …). Кастомные пользовательские таблицы — roadmap.
| `extract_mode` | Что делает | Откуда в UI |
|----------------|------------|-------------|
| `heuristic` | Legacy regex-парсер постов | raw job API (не пара) |
| `llm` | DeepSeek **на каждый** пост (batch) | профиль `kind=llm` → pair flatten |
| `profile` | Статичные правила `HeuristicProfile` | профиль `kind=heuristic` → pair flatten |
### LLM-режим (`extract_mode: llm`)
Для неструктурированных Telegram-постов и Crawl4AI:
- ключ `DEEPSEEK_API_KEY` в `.env` (воркеры `cp-workers` / `cp-workers-web`);
- Telegram: текст поста → DeepSeek JSON → `IngestEventItem`;
- ключ `DEEPSEEK_API_KEY` в `.env` (воркеры `cp-workers` / `cp-workers-web`; preview в `ca-api`);
- Telegram batch: текст → DeepSeek JSON → один или несколько `IngestEventItem` (`workers/llm_extract.py`);
- опционально `extract_schema`, `instruction`, `required_fields` (post-extract gate поверх `is_event`);
- **`multi_event`** (opt-in в `llm_profile` / `TelegramSourceConfig`): модель возвращает массив `events`; каждое событие — отдельный ingest. URL: один event → `post.url`; несколько → `post.url#e1`, `#e2`, … (дедуп ЦА по `source_url`);
- heuristic / profile по-прежнему **1 пост → 1 событие**;
- Crawl4AI: страница → `LLMExtractionStrategy` (DeepSeek) с fallback на тот же DeepSeek по markdown;
- в UI «Парсеры»: поле **Извлечение** = LLM DeepSeek.
- **listener:** для `llm` пока fallback на heuristic (LLM — batch-only by design).
Reusable LLM-профиль в ЦА: `ParserProfile.kind=llm` + JSON `llm_profile` (`contracts/llm_profile.py`). При enqueue flatten → `extract_mode=llm` + schema/instruction/required_fields/`multi_event`.
### Profile-режим (`extract_mode: profile`)
@@ -69,8 +84,12 @@ async def run(job_id, source_config, *, ctx) -> tuple[list[dict], str | None]
- `heuristic_profile` в `source_config` (схема `contracts/heuristic_profile.py`);
- интерпретатор: `workers/heuristic_profile.py` (те же правила, что preview в ЦА);
- Telegram batch + listener применяют профиль без DeepSeek;
- генерация профиля — только в админке (`/admin/parser-profiles/generate`).
- `required_fields`: пост без заполненных обязательных полей не ингестится (batch + listener);
- генерация правил — один раз в админке (`/admin/parser-profiles/generate`); runtime **без** LLM.
Reusable heuristic-профиль: `ParserProfile.kind=heuristic` + `heuristic_profile`. Pair на «Парсеры» → flatten `extract_mode=profile`.
UI «Профили»: выбор `kind` (heuristic | llm). Связка канал+профиль на «Парсеры».
```mermaid
flowchart LR
@@ -127,7 +146,10 @@ Legacy-ключ `cp:jobs` по-прежнему дренируется telegram-
## Telegram real-time
`TelegramListener` без изменений: подписки из ЦА, ingest с `listener: true` (статус `ParseJob` не трогается).
`TelegramListener`: подписки из ЦА, ingest с `listener: true` (статус `ParseJob` не трогается).
- `extract_mode=profile` — те же heuristic-правила, что batch;
- `extract_mode=llm` — **не** вызывает DeepSeek; fallback на legacy heuristic (LLM только в batch).
## Переменные окружения
@@ -11,9 +11,11 @@ from workers.heuristic_profile import extract_with_profile
from workers.llm_extract import (
DEFAULT_EXTRACT_SCHEMA,
DEFAULT_INSTRUCTION,
extract_event_fields,
event_source_url,
extract_event_list,
fields_to_ingest,
llm_enabled,
match_llm_required,
)
from workers.parsers.telegram_events import parse_event_posts
from workers.sources.telegram_client import (
@@ -73,54 +75,72 @@ def _extract_posts_with_profile(posts, cfg: TelegramSourceConfig) -> tuple[list[
text = (post.text or "").strip()
if not text:
continue
events.append(
extract_with_profile(
text,
profile,
source_url=post.url,
source_type="telegram",
extra_metadata={
"channel": post.channel,
"message_id": post.id,
"post_date": post.date.isoformat() if post.date else None,
},
)
event = extract_with_profile(
text,
profile,
source_url=post.url,
source_type="telegram",
extra_metadata={
"channel": post.channel,
"message_id": post.id,
"post_date": post.date.isoformat() if post.date else None,
},
)
if event is None:
logger.debug("Profile skip %s (required fields missing)", post.url)
continue
events.append(event)
return events, None
async def _extract_posts_with_llm(posts, cfg: TelegramSourceConfig) -> tuple[list[dict], str | None]:
schema = cfg.extract_schema or DEFAULT_EXTRACT_SCHEMA
instruction = cfg.instruction or DEFAULT_INSTRUCTION
multi_event = bool(cfg.multi_event)
events: list[dict] = []
errors: list[str] = []
required = list(cfg.required_fields or [])
for post in posts:
text = (post.text or "").strip()
if not text:
continue
try:
fields = await extract_event_fields(
field_list = await extract_event_list(
text,
extract_schema=schema,
instruction=instruction,
multi_event=multi_event,
)
if not fields.get("is_event", True):
if not field_list:
continue
events.append(
fields_to_ingest(
source_type="telegram",
source_url=post.url,
raw_text=text,
fields=fields,
domain_profile="telegram_llm",
extra_metadata={
"channel": post.channel,
"message_id": post.id,
"post_date": post.date.isoformat() if post.date else None,
},
total = len(field_list)
for idx, fields in enumerate(field_list):
if not match_llm_required(fields, required):
logger.debug(
"LLM skip %s event %s/%s (required fields missing)",
post.url,
idx + 1,
total,
)
continue
events.append(
fields_to_ingest(
source_type="telegram",
source_url=event_source_url(post.url, idx, total),
raw_text=text,
fields=fields,
domain_profile="telegram_llm",
extra_metadata={
"channel": post.channel,
"message_id": post.id,
"post_date": post.date.isoformat() if post.date else None,
"event_index": idx + 1,
"event_count": total,
"multi_event": multi_event and total > 1,
},
)
)
)
except Exception as exc:
logger.exception("LLM extract failed for %s", post.url)
errors.append(f"{post.url}: {exc}")
@@ -4,7 +4,7 @@ from __future__ import annotations
from typing import Any
from contracts.heuristic_profile import HeuristicProfile, apply_profile
from contracts.heuristic_profile import HeuristicProfile, match_profile
from workers.llm_extract import parse_coords, parse_date
@@ -62,8 +62,10 @@ def extract_with_profile(
source_url: str,
source_type: str = "telegram",
extra_metadata: dict | None = None,
) -> dict:
fields = apply_profile(text, profile)
) -> dict | None:
matched, fields, _missing = match_profile(text, profile)
if not matched:
return None
return profile_fields_to_ingest(
source_type=source_type,
source_url=source_url,
+128 -48
View File
@@ -5,30 +5,31 @@ from __future__ import annotations
import json
import logging
import os
import re
from datetime import datetime, timezone
from typing import Any
import httpx
from contracts.heuristic_profile import parse_coords
from contracts.llm_profile import (
DEFAULT_EXTRACT_SCHEMA,
DEFAULT_INSTRUCTION,
match_llm_required,
)
logger = logging.getLogger("cp-worker.llm")
COORDS_RE = re.compile(r"(-?\d{1,3}\.\d+)\s*,\s*(-?\d{1,3}\.\d+)")
DEFAULT_EXTRACT_SCHEMA: dict[str, str] = {
"title": "string — short event title",
"locality": "string — place / settlement name",
"event_date": "string — date as DD.MM.YYYY or YYYY-MM-DD if known",
"description": "string — concise event summary",
"coords": "string — latitude, longitude if present else empty",
"topic": "string — short topic tag",
}
DEFAULT_INSTRUCTION = (
"Extract structured military/news event fields from the text. "
"If the text is not an event, return is_event=false. "
"Respond with a single JSON object only."
)
# Re-export for adapters that import from this module
__all__ = [
"DEFAULT_EXTRACT_SCHEMA",
"DEFAULT_INSTRUCTION",
"event_source_url",
"extract_event_fields",
"extract_event_list",
"fields_to_ingest",
"llm_enabled",
"match_llm_required",
]
def llm_enabled() -> bool:
@@ -43,29 +44,57 @@ def llm_settings() -> dict[str, str]:
}
async def extract_event_fields(
text: str,
def event_source_url(post_url: str, index: int, total: int) -> str:
"""Stable per-event URL for CA dedup. Single event keeps bare post URL."""
base = (post_url or "").strip()
if total <= 1:
return base
# index is 0-based; fragment uses 1-based #eN
return f"{base}#e{index + 1}"
def _normalize_fields(raw: dict[str, Any], schema: dict[str, str]) -> dict[str, str]:
return {key: str(raw.get(key) or "").strip() for key in schema}
def _parse_llm_content(
content: str,
schema: dict[str, str],
*,
extract_schema: dict[str, str] | None = None,
instruction: str | None = None,
) -> dict[str, Any]:
"""Ask DeepSeek to fill schema fields from free text. Returns dict (+ is_event)."""
multi_event: bool,
) -> tuple[bool, list[dict[str, str]]]:
"""Return (is_event, list of field dicts)."""
parsed = json.loads(content)
if not isinstance(parsed, dict):
return False, []
is_event = bool(parsed.get("is_event", True))
if not is_event:
return False, []
if multi_event and isinstance(parsed.get("events"), list):
events: list[dict[str, str]] = []
for item in parsed["events"]:
if isinstance(item, dict):
events.append(_normalize_fields(item, schema))
return True, events
# Single-event fallback (legacy fields / top-level keys)
fields_raw = parsed.get("fields") if isinstance(parsed.get("fields"), dict) else parsed
if not isinstance(fields_raw, dict):
fields_raw = {}
# Drop non-schema keys that confuse normalize when falling back to top-level
cleaned = {k: fields_raw.get(k) for k in schema}
return True, [_normalize_fields(cleaned, schema)]
async def _call_deepseek(user_prompt: str) -> str:
settings = llm_settings()
if not settings["api_key"]:
raise RuntimeError(
"DEEPSEEK_API_KEY is not set. Add it to .env for LLM extract_mode."
)
schema = extract_schema or DEFAULT_EXTRACT_SCHEMA
instr = instruction or DEFAULT_INSTRUCTION
schema_lines = "\n".join(f"- {k}: {v}" for k, v in schema.items())
user_prompt = (
f"{instr}\n\n"
f"Fields to extract:\n{schema_lines}\n\n"
'Return JSON: {"is_event": true|false, "fields": {<field>: <string>}}\n\n'
f"Text:\n{text[:12000]}"
)
payload = {
"model": settings["model"],
"messages": [
@@ -95,14 +124,72 @@ async def extract_event_fields(
response.raise_for_status()
data = response.json()
content = data["choices"][0]["message"]["content"]
parsed = json.loads(content)
fields = parsed.get("fields") if isinstance(parsed.get("fields"), dict) else parsed
if not isinstance(fields, dict):
fields = {}
# Normalize to strings for known keys
result = {key: str(fields.get(key) or "").strip() for key in schema}
result["is_event"] = bool(parsed.get("is_event", True))
return data["choices"][0]["message"]["content"]
async def extract_event_list(
text: str,
*,
extract_schema: dict[str, str] | None = None,
instruction: str | None = None,
multi_event: bool = False,
) -> list[dict[str, Any]]:
"""Extract zero or more event field dicts from free text.
Each dict has schema keys as strings. Empty list if not an event / no items.
"""
schema = extract_schema or DEFAULT_EXTRACT_SCHEMA
instr = instruction or DEFAULT_INSTRUCTION
schema_lines = "\n".join(f"- {k}: {v}" for k, v in schema.items())
if multi_event:
user_prompt = (
f"{instr}\n\n"
"If the text describes multiple distinct events (different places, "
"coords, or dates), return one object per event in \"events\".\n"
f"Fields per event:\n{schema_lines}\n\n"
'Return JSON: {"is_event": true|false, "events": [{<field>: <string>}, ...]}\n'
"If there is no event, return is_event=false and events=[].\n\n"
f"Text:\n{text[:12000]}"
)
else:
user_prompt = (
f"{instr}\n\n"
f"Fields to extract:\n{schema_lines}\n\n"
'Return JSON: {"is_event": true|false, "fields": {<field>: <string>}}\n\n'
f"Text:\n{text[:12000]}"
)
content = await _call_deepseek(user_prompt)
is_event, events = _parse_llm_content(content, schema, multi_event=multi_event)
if not is_event:
return []
return events
async def extract_event_fields(
text: str,
*,
extract_schema: dict[str, str] | None = None,
instruction: str | None = None,
) -> dict[str, Any]:
"""Ask DeepSeek to fill schema fields from free text. Returns dict (+ is_event).
Single-event API kept for crawl4ai and legacy callers.
"""
events = await extract_event_list(
text,
extract_schema=extract_schema,
instruction=instruction,
multi_event=False,
)
if not events:
schema = extract_schema or DEFAULT_EXTRACT_SCHEMA
empty = {key: "" for key in schema}
empty["is_event"] = False
return empty
result = dict(events[0])
result["is_event"] = True
return result
@@ -146,13 +233,6 @@ def fields_to_ingest(
}
def parse_coords(raw: str) -> tuple[float | None, float | None]:
match = COORDS_RE.search(raw or "")
if not match:
return None, None
return float(match.group(1)), float(match.group(2))
def parse_date(raw: str) -> datetime | None:
if not raw:
return None
@@ -81,7 +81,7 @@ class TelegramListener:
", ".join(sorted(channels)) or "(none)",
)
def _build_event(self, channel: str, post) -> dict:
def _build_event(self, channel: str, post) -> dict | None:
sub = self._channels.get(normalize_channel(channel)) or {}
cfg = sub.get("source_config") or {}
extract_mode = cfg.get("extract_mode") or "heuristic"
@@ -107,6 +107,9 @@ class TelegramListener:
async def _ingest_post(self, channel: str, post) -> None:
event = self._build_event(channel, post)
if event is None:
logger.debug("Listener skip %s (profile required fields missing)", post.url)
return
sub = self._channels.get(normalize_channel(channel)) or {}
job_id = sub.get("job_id")
+53
View File
@@ -17,6 +17,8 @@ TARGET_FIELDS: tuple[str, ...] = (
"region",
)
COORDS_RE = re.compile(r"(-?\d{1,3}\.\d+)\s*,\s*(-?\d{1,3}\.\d+)")
Strategy = Literal["regex", "line", "after_marker", "between", "full_text", "literal"]
@@ -45,8 +47,24 @@ class FieldRule(BaseModel):
class HeuristicProfile(BaseModel):
version: Literal[1] = 1
fields: dict[str, FieldRule] = Field(default_factory=dict)
required_fields: list[str] = Field(default_factory=list)
notes: str = ""
@field_validator("required_fields")
@classmethod
def known_required(cls, value: list[str]) -> list[str]:
seen: list[str] = []
unknown: list[str] = []
for name in value:
if name not in TARGET_FIELDS:
unknown.append(name)
continue
if name not in seen:
seen.append(name)
if unknown:
raise ValueError(f"Unknown required_fields: {sorted(set(unknown))}")
return seen
@model_validator(mode="after")
def known_fields_only(self) -> "HeuristicProfile":
unknown = set(self.fields) - set(TARGET_FIELDS)
@@ -163,3 +181,38 @@ def apply_profile(text: str, profile: HeuristicProfile | dict[str, Any]) -> dict
for name, rule in profile.fields.items():
result[name] = _apply_rule(text, rule)
return result
def parse_coords(raw: str) -> tuple[float | None, float | None]:
match = COORDS_RE.search(raw or "")
if not match:
return None, None
return float(match.group(1)), float(match.group(2))
def _field_is_filled(name: str, value: str) -> bool:
if not (value or "").strip():
return False
if name == "coords":
lat, lng = parse_coords(value)
return lat is not None and lng is not None
return True
def match_profile(
text: str,
profile: HeuristicProfile | dict[str, Any],
) -> tuple[bool, dict[str, str], list[str]]:
"""Apply profile and report whether required_fields are filled.
Empty required_fields means no gate (legacy profiles ingest every post).
"""
if isinstance(profile, dict):
profile = HeuristicProfile.model_validate(profile)
fields = apply_profile(text, profile)
missing = [
name
for name in profile.required_fields
if not _field_is_filled(name, fields.get(name, ""))
]
return (not missing, fields, missing)
+86
View File
@@ -0,0 +1,86 @@
"""Reusable LLM extract profile (CA storage + flattened into TelegramSourceConfig)."""
from __future__ import annotations
from typing import Any
from pydantic import BaseModel, Field, field_validator, model_validator
from contracts.heuristic_profile import TARGET_FIELDS
# LLM defaults omit region (same as workers/llm_extract.DEFAULT_EXTRACT_SCHEMA).
LLM_SCHEMA_FIELDS: tuple[str, ...] = (
"title",
"locality",
"event_date",
"description",
"coords",
"topic",
)
DEFAULT_EXTRACT_SCHEMA: dict[str, str] = {
"title": "string — short event title",
"locality": "string — place / settlement name",
"event_date": "string — date as DD.MM.YYYY or YYYY-MM-DD if known",
"description": "string — concise event summary",
"coords": "string — latitude, longitude if present else empty",
"topic": "string — short topic tag",
}
DEFAULT_INSTRUCTION = (
"Extract structured military/news event fields from the text. "
"If the text is not an event, return is_event=false. "
"Respond with a single JSON object only."
)
class LlmProfile(BaseModel):
instruction: str | None = None
extract_schema: dict[str, str] = Field(default_factory=lambda: dict(DEFAULT_EXTRACT_SCHEMA))
required_fields: list[str] = Field(default_factory=list)
# When true, LLM may return multiple events per post (array "events")
multi_event: bool = False
@field_validator("instruction", mode="before")
@classmethod
def empty_instruction_to_none(cls, value: Any) -> Any:
if value is None:
return None
if isinstance(value, str) and not value.strip():
return None
return value
@field_validator("required_fields")
@classmethod
def known_required(cls, value: list[str]) -> list[str]:
seen: list[str] = []
unknown: list[str] = []
for name in value:
if name not in TARGET_FIELDS:
unknown.append(name)
continue
if name not in seen:
seen.append(name)
if unknown:
raise ValueError(f"Unknown required_fields: {sorted(set(unknown))}")
return seen
@model_validator(mode="after")
def known_schema_keys(self) -> "LlmProfile":
if not self.extract_schema:
raise ValueError("extract_schema must not be empty")
unknown = set(self.extract_schema) - set(TARGET_FIELDS)
if unknown:
raise ValueError(f"Unknown extract_schema fields: {sorted(unknown)}")
return self
def match_llm_required(fields: dict[str, Any], required_fields: list[str]) -> bool:
"""True if all required_fields are non-empty strings (empty required = no gate)."""
if not required_fields:
return True
for name in required_fields:
val = fields.get(name)
if val is None or not str(val).strip():
return False
return True
+34 -8
View File
@@ -15,6 +15,10 @@ class TelegramSourceConfig(BaseModel):
extract_schema: dict[str, str] | None = None
instruction: str | None = None
heuristic_profile: dict | None = None
# Post-extract gate for extract_mode=llm (empty = only is_event filter)
required_fields: list[str] = Field(default_factory=list)
# Split one post into N events when extract_mode=llm (opt-in)
multi_event: bool = False
sample_post: str | None = None # audit / re-generate; not required at runtime
@field_validator("channel")
@@ -22,15 +26,37 @@ class TelegramSourceConfig(BaseModel):
def strip_channel(cls, value: str) -> str:
return value.strip()
@model_validator(mode="after")
def profile_requires_rules(self) -> "TelegramSourceConfig":
if self.extract_mode != "profile":
return self
if not self.heuristic_profile:
raise ValueError("heuristic_profile required when extract_mode=profile")
from contracts.heuristic_profile import HeuristicProfile
@field_validator("required_fields")
@classmethod
def known_required(cls, value: list[str]) -> list[str]:
from contracts.heuristic_profile import TARGET_FIELDS
HeuristicProfile.model_validate(self.heuristic_profile)
seen: list[str] = []
unknown: list[str] = []
for name in value:
if name not in TARGET_FIELDS:
unknown.append(name)
continue
if name not in seen:
seen.append(name)
if unknown:
raise ValueError(f"Unknown required_fields: {sorted(set(unknown))}")
return seen
@model_validator(mode="after")
def validate_extract_mode(self) -> "TelegramSourceConfig":
if self.extract_mode == "profile":
if not self.heuristic_profile:
raise ValueError("heuristic_profile required when extract_mode=profile")
from contracts.heuristic_profile import HeuristicProfile
HeuristicProfile.model_validate(self.heuristic_profile)
elif self.extract_mode == "llm" and self.extract_schema is not None:
from contracts.heuristic_profile import TARGET_FIELDS
unknown = set(self.extract_schema) - set(TARGET_FIELDS)
if unknown:
raise ValueError(f"Unknown extract_schema fields: {sorted(unknown)}")
return self
+7 -1
View File
@@ -184,7 +184,12 @@ async def run(job_id, source_config, *, ctx) -> tuple[list[dict], str | None]
Опционально: `extract_mode: llm` (DeepSeek) для telegram/crawl4ai в batch.
**Профиль + канал:** UI `/parser-profiles` (Generate/Preview/CRUD) и `/channels`; связка на `/parsers` → `ParseJob` с FK. При enqueue ЦА flatten в `extract_mode=profile` + `heuristic_profile`. Runtime (batch + listener) применяет только профиль, без LLM. Кастомные пользовательские таблицы — roadmap.
**Профиль + канал:** UI `/parser-profiles` (CRUD, Generate/Preview для heuristic; instruction/schema/Preview для llm) и `/channels`; связка на `/parsers` → `ParseJob` с FK. При enqueue ЦА flatten по `ParserProfile.kind`:
- `heuristic` → `extract_mode=profile` + `heuristic_profile` (runtime без LLM; listener поддерживает);
- `llm` → `extract_mode=llm` + `extract_schema` / `instruction` / `required_fields` (DeepSeek на каждый пост в batch; listener пока fallback на heuristic).
Целевые поля — фиксированная схема Event/`IngestEventItem`. Кастомные пользовательские таблицы — roadmap.
Подробности: [centers/parsing/ARCHITECTURE.md](../centers/parsing/ARCHITECTURE.md), [centers/analytics/ARCHITECTURE.md](../centers/analytics/ARCHITECTURE.md).
@@ -209,6 +214,7 @@ async def run(job_id, source_config, *, ctx) -> tuple[list[dict], str | None]
| `contracts/sources.py` | Валидация `source_config` по `source_type` |
| `contracts/queues.py` | `SOURCE_FAMILY` → ключ очереди |
| `contracts/heuristic_profile.py` | Статичный профиль конструктора + `apply_profile` |
| `contracts/llm_profile.py` | Reusable LLM instruction/schema + `match_llm_required` |
Правило: меняете форму события / конфиг источника / очередь — сначала `contracts/`, потом CA/CP/UI.
+12 -3
View File
@@ -13,7 +13,9 @@
| [`jobs.py`](../contracts/jobs.py) | Контракт задания. Payload задания в Redis |
| [`ingest.py`](../contracts/ingest.py) |Контракт результата. Событие и пакет ingest |
| [`sources.py`](../contracts/sources.py) | Контракт настроек парсера. Схемы `source_config` по `source_type` |
| [`queues.py`](../contracts/queues.py) | Контракт доставки задания нужному воркеру. `source_type` → family → Redis key |
| [`queues.py`](../contracts/queues.py) | Контракт доставки задания нужному воркеру. `source_type` → family → Redis key |
| [`heuristic_profile.py`](../contracts/heuristic_profile.py) | Статичные правила extract_mode=profile |
| [`llm_profile.py`](../contracts/llm_profile.py) | Reusable LLM instruction/schema + required_fields + `multi_event` |
ЦА при admin CRUD валидирует конфиг через `parse_source_config`.
ЦП адаптеры должны отдавать dict, совместимые с `IngestEventItem`.
@@ -57,13 +59,20 @@ Listener добавляет флаг `listener: true` на стороне CA API
| source_type | Модель | Главные поля |
|-------------|--------|--------------|
| `telegram` | `TelegramSourceConfig` | `channel`, `limit`, `extract_mode` (`heuristic`\|`llm`\|`profile`), `heuristic_profile` |
| `telegram` | `TelegramSourceConfig` | `channel`, `limit`, `extract_mode` (`heuristic`\|`llm`\|`profile`), `heuristic_profile`, `extract_schema`, `instruction`, `required_fields` (llm gate) |
| `crawl4ai` | `Crawl4AISourceConfig` | `urls`, `extract_mode`, `extract_schema`, `domain_profile` |
| `viina` | `ViinaSourceConfig` | `urls` / `texts`, `input_mode` |
Реестр: `CONFIG_MODELS` + `parse_source_config(source_type, raw)`.
Профиль парсера: `contracts/heuristic_profile.py` (`HeuristicProfile`, `apply_profile`). Сущности Profile/Channel живут в БД ЦА; в Redis уходит плоский `heuristic_profile`.
Профили в БД ЦА (`ParserProfile`):
| kind | Хранилище | Flatten в Redis |
|------|-----------|-----------------|
| `heuristic` | `heuristic_profile` (`contracts/heuristic_profile.py`) | `extract_mode=profile` |
| `llm` | `llm_profile` (`contracts/llm_profile.py`: instruction, extract_schema, required_fields, multi_event) | `extract_mode=llm` (+ `#eN` URLs if multi) |
`required_fields` — обязательные поля; пустой список = без фильтра (для llm остаётся только `is_event`). Канал + профиль живут в ЦА; CP получает только плоский `source_config`.
---