Add parser builder: one-shot DeepSeek profile for Telegram extract_mode=profile.

Generate static HeuristicProfile in CA admin, preview and run without LLM on each post via shared interpreter in CP batch and listener.

Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
2026-08-16 14:14:18 +03:00
co-authored by Cursor
parent 6dff3c1c3d
commit 699e9be503
23 changed files with 1042 additions and 20 deletions
+3 -1
View File
@@ -12,7 +12,9 @@ TELEGRAM_SESSION_PATH=/data/telegram.session
INTERNAL_TOKEN=dev-internal-token
TEST_PI_API_KEY=test-pi-api-key-change-me
# DeepSeek LLM for extract_mode=llm (Telegram unstructured + Crawl4AI LLM)
# DeepSeek LLM:
# - extract_mode=llm on CP workers (Telegram unstructured + Crawl4AI)
# - parser-builder Generate on ca-api (one-shot profile; runtime stays rule-based)
# DEEPSEEK_API_KEY=sk-...
# DEEPSEEK_BASE_URL=https://api.deepseek.com
# DEEPSEEK_MODEL=deepseek-chat
+18 -2
View File
@@ -59,12 +59,14 @@ centers/analytics/
│ │ ├── objects.py # /api/health, /api/objects, media
│ │ ├── map.py # /api/map/*
│ │ ├── admin.py # /admin/*
│ │ ├── parser_builder.py # /admin/parser-builder/*
│ │ ├── internal.py # /internal/* (только ЦП)
│ │ └── v1.py # /api/v1/* (ПИ)
│ └── services/
│ ├── jobs.py # Redis RPUSH
│ ├── scheduler.py # периодический re-queue
│ ├── ingest.py # дедуп + map sync
│ ├── parser_builder.py # DeepSeek → HeuristicProfile (один раз)
│ ├── filtering.py
│ ├── map_query.py
│ ├── events_query.py
@@ -74,7 +76,7 @@ centers/analytics/
├── Dockerfile # Vite build + nginx
├── nginx.conf # proxy /api /admin /internal → ca-api
└── src/
├── views/ # Map, Parsers, Events, Analytics, Consumers
├── views/ # Map, Parsers, ParserBuilder, Events, …
├── components/ # карта, CRUD объектов
├── api/ # HTTP-клиенты
└── router/index.ts
@@ -96,12 +98,23 @@ centers/analytics/
| Prefix | Кто вызывает | Содержание |
|--------|--------------|------------|
| `/api/*` | UI, публичный health | Карта, объекты, медиа |
| `/admin/*` | UI admin | Jobs, events, analytics, consumers |
| `/admin/*` | UI admin | Jobs, events, analytics, consumers, parser-builder |
| `/internal/*` | Только ЦП | ingest, job status, listener subscriptions |
| `/api/v1/*` | Внешние клиенты | Events с Bearer-ключом |
Internal защищён заголовком `X-Internal-Token` (`INTERNAL_TOKEN`).
### Конструктор парсера
Раздел UI `/parser-builder` → `POST /admin/parser-builder/generate|preview|jobs`:
1. Менеджер вставляет образец поста.
2. DeepSeek (`DEEPSEEK_API_KEY` на **ca-api**) один раз возвращает `HeuristicProfile` (regex/line/marker).
3. Preview и сохранение job с `extract_mode=profile` + JSON профиля в `source_config`.
4. ЦП применяет профиль статически (batch + listener); LLM на ingest не вызывается.
Roadmap: кастомные пользовательские таблицы подменяют только список `target-fields` при генерации.
### Jobs и scheduler
1. `POST /admin/jobs` / retry → запись `ParseJob` + `enqueue_job` (`services/jobs.py`).
@@ -123,6 +136,7 @@ Internal защищён заголовком `X-Internal-Token` (`INTERNAL_TOKEN
|------|------|
| `/` | `MapViewPage.vue` |
| `/parsers` | `ParsersView.vue` |
| `/parser-builder` | `ParserBuilderView.vue` |
| `/events` | `EventsView.vue` |
| `/analytics` | `AnalyticsView.vue` |
| `/consumers` | `ConsumersView.vue` |
@@ -142,6 +156,8 @@ Internal защищён заголовком `X-Internal-Token` (`INTERNAL_TOKEN
| `INTERNAL_TOKEN` | Auth ЦП ↔ ЦА |
| `TEST_PI_API_KEY` | Seed consumer `test-pi` |
| `UPLOAD_DIR` | Медиа (по умолчанию `/data/uploads`) |
| `DEEPSEEK_API_KEY` | Конструктор парсера (Generate); не нужен для preview/runtime profile |
| `DEEPSEEK_BASE_URL` / `DEEPSEEK_MODEL` | Опционально |
### Типовые точки входа в код
+2 -1
View File
@@ -13,7 +13,7 @@ for _candidate in (_HERE.parent, *_HERE.parents):
break
from .database import Base, engine, get_db
from .routers import admin, internal, map, objects, v1
from .routers import admin, internal, map, objects, parser_builder, v1
from .seed import seed_objects, seed_test_consumer
from .services.migrations import migrate_schema
from .services.scheduler import start_scheduler
@@ -50,4 +50,5 @@ app.include_router(objects.router)
app.include_router(map.router)
app.include_router(internal.router)
app.include_router(admin.router)
app.include_router(parser_builder.router)
app.include_router(v1.router)
@@ -71,5 +71,11 @@ def listener_subscriptions(
channel = job.source_config.get("channel") if job.source_config else None
if not channel:
continue
result.append(ListenerSubscription(job_id=job.id, channel=str(channel)))
result.append(
ListenerSubscription(
job_id=job.id,
channel=str(channel),
source_config=dict(job.source_config or {}),
)
)
return result
@@ -0,0 +1,108 @@
"""Admin API: construct static telegram parsers from a sample post."""
from __future__ import annotations
from typing import Any
from fastapi import APIRouter, Depends, HTTPException
from pydantic import BaseModel, Field
from sqlalchemy.orm import Session
from contracts.heuristic_profile import HeuristicProfile
from contracts.sources import TelegramSourceConfig
from ..database import get_db
from ..models import ParseJob
from ..schemas import ParseJobRead
from ..services.jobs import enqueue_job
from ..services import parser_builder as builder
router = APIRouter(prefix="/admin/parser-builder", tags=["parser-builder"])
class GenerateRequest(BaseModel):
sample_post: str = Field(min_length=1)
class GenerateResponse(BaseModel):
profile: dict[str, Any]
class PreviewRequest(BaseModel):
sample_post: str = Field(min_length=1)
profile: dict[str, Any]
class PreviewResponse(BaseModel):
fields: dict[str, str]
class CreateProfileJobRequest(BaseModel):
channel: str = Field(min_length=1)
profile: dict[str, Any]
sample_post: str | None = None
limit: int = Field(default=100, ge=1, le=1000)
interval_seconds: int = Field(default=3600, ge=60, le=604800)
is_active: bool = True
@router.get("/target-fields")
def target_fields():
return {"fields": builder.get_target_fields()}
@router.post("/generate", response_model=GenerateResponse)
async def generate_profile(payload: GenerateRequest):
if not builder.deepseek_enabled():
raise HTTPException(
status_code=503,
detail="DEEPSEEK_API_KEY is not set on ca-api. Add it to .env for parser generation.",
)
try:
profile = await builder.generate_profile(payload.sample_post)
except ValueError as exc:
raise HTTPException(status_code=400, detail=str(exc)) from exc
except Exception as exc:
raise HTTPException(
status_code=502,
detail=f"DeepSeek generate failed: {exc}",
) from exc
return GenerateResponse(profile=profile.model_dump())
@router.post("/preview", response_model=PreviewResponse)
def preview_profile(payload: PreviewRequest):
try:
HeuristicProfile.model_validate(payload.profile)
fields = builder.preview_with_profile(payload.sample_post, payload.profile)
except Exception as exc:
raise HTTPException(status_code=400, detail=str(exc)) from exc
return PreviewResponse(fields=fields)
@router.post("/jobs", response_model=ParseJobRead, status_code=201)
def create_profile_job(payload: CreateProfileJobRequest, db: Session = Depends(get_db)):
try:
cfg = TelegramSourceConfig(
channel=payload.channel.strip(),
limit=payload.limit,
extract_mode="profile",
heuristic_profile=payload.profile,
sample_post=payload.sample_post,
)
except Exception as exc:
raise HTTPException(status_code=400, detail=str(exc)) from exc
job = ParseJob(
source_type="telegram",
source_config=cfg.model_dump(),
interval_seconds=payload.interval_seconds,
is_active=payload.is_active,
status="queued",
)
db.add(job)
db.commit()
db.refresh(job)
enqueue_job(job.id, job.source_type, job.source_config)
return job
+1
View File
@@ -117,6 +117,7 @@ class IngestRequest(BaseModel):
class ListenerSubscription(BaseModel):
job_id: int
channel: str
source_config: dict[str, Any] = Field(default_factory=dict)
class IngestResponse(BaseModel):
@@ -0,0 +1,130 @@
"""Parser builder: one-shot DeepSeek profile generation + static preview."""
from __future__ import annotations
import json
import logging
import os
import re
from typing import Any
import httpx
from contracts.heuristic_profile import (
HeuristicProfile,
TARGET_FIELDS,
apply_profile,
target_field_specs,
)
logger = logging.getLogger("ca.parser_builder")
GENERATE_SYSTEM = (
"You design a static heuristic parser profile for structured Telegram posts. "
"Output valid JSON only, no markdown, no Python code. "
"The profile is applied with regex/line/marker rules at runtime — never with an LLM."
)
def deepseek_enabled() -> bool:
return bool(os.getenv("DEEPSEEK_API_KEY", "").strip())
def deepseek_settings() -> dict[str, str]:
return {
"api_key": os.getenv("DEEPSEEK_API_KEY", "").strip(),
"base_url": os.getenv("DEEPSEEK_BASE_URL", "https://api.deepseek.com").rstrip("/"),
"model": os.getenv("DEEPSEEK_MODEL", "deepseek-chat"),
}
def get_target_fields() -> list[dict[str, str]]:
return target_field_specs()
def preview_with_profile(sample_post: str, profile: dict[str, Any] | HeuristicProfile) -> dict[str, str]:
return apply_profile(sample_post, profile)
def _extract_json_object(content: str) -> dict[str, Any]:
content = (content or "").strip()
if content.startswith("```"):
content = re.sub(r"^```(?:json)?\s*", "", content)
content = re.sub(r"\s*```$", "", content)
try:
parsed = json.loads(content)
except json.JSONDecodeError:
match = re.search(r"\{[\s\S]*\}", content)
if not match:
raise
parsed = json.loads(match.group(0))
if not isinstance(parsed, dict):
raise ValueError("LLM response must be a JSON object")
return parsed
async def generate_profile(sample_post: str) -> HeuristicProfile:
settings = deepseek_settings()
if not settings["api_key"]:
raise RuntimeError(
"DEEPSEEK_API_KEY is not set on ca-api. Add it to .env for parser generation."
)
sample = (sample_post or "").strip()
if not sample:
raise ValueError("sample_post is required")
fields_help = "\n".join(
f"- {spec['name']}: {spec['description']}" for spec in target_field_specs()
)
user_prompt = (
"Build a HeuristicProfile JSON for this sample Telegram post.\n\n"
"Schema:\n"
'{"version": 1, "notes": "...", "fields": {'
'"<field>": {"strategy": "regex|line|after_marker|between|full_text|literal", '
'"pattern": "...", "group": 1, "line_index": 0, "marker": "...", '
'"end_marker": "...", "value": "...", "flags": "im", "strip": true}'
"}}\n\n"
"Rules:\n"
"- Only include fields you can extract reliably from the sample.\n"
f"- Allowed field names: {', '.join(TARGET_FIELDS)}.\n"
"- Prefer regex / line / after_marker / between over literal.\n"
"- Use literal only for constant topic tags.\n"
"- coords should capture 'lat, lon' when present.\n"
"- event_date should capture DD.MM.YYYY or similar.\n"
"- Do not invent Python code; only declarative rules.\n\n"
f"Target fields:\n{fields_help}\n\n"
f"Sample post:\n{sample[:12000]}"
)
payload = {
"model": settings["model"],
"messages": [
{"role": "system", "content": GENERATE_SYSTEM},
{"role": "user", "content": user_prompt},
],
"temperature": 0.1,
"response_format": {"type": "json_object"},
}
url = f"{settings['base_url']}/chat/completions"
async with httpx.AsyncClient(timeout=90.0) as client:
response = await client.post(
url,
headers={
"Authorization": f"Bearer {settings['api_key']}",
"Content-Type": "application/json",
},
json=payload,
)
response.raise_for_status()
data = response.json()
content = data["choices"][0]["message"]["content"]
raw = _extract_json_object(content)
# Accept either top-level profile or {"profile": {...}}
if "fields" not in raw and isinstance(raw.get("profile"), dict):
raw = raw["profile"]
if "version" not in raw:
raw["version"] = 1
return HeuristicProfile.model_validate(raw)
+1
View File
@@ -5,3 +5,4 @@ pydantic==2.10.3
python-multipart==0.0.20
psycopg2-binary==2.9.10
redis==5.2.1
httpx==0.28.1
@@ -0,0 +1,52 @@
import { ADMIN_BASE, request } from "./client";
import type { ParseJob } from "../types/admin";
export type TargetField = {
name: string;
type: string;
description: string;
};
export type HeuristicProfile = {
version: number;
fields: Record<string, Record<string, unknown>>;
notes?: string;
};
export function fetchTargetFields(): Promise<{ fields: TargetField[] }> {
return request<{ fields: TargetField[] }>("/parser-builder/target-fields", undefined, ADMIN_BASE);
}
export function generateParserProfile(sample_post: string): Promise<{ profile: HeuristicProfile }> {
return request<{ profile: HeuristicProfile }>(
"/parser-builder/generate",
{ method: "POST", body: JSON.stringify({ sample_post }) },
ADMIN_BASE,
);
}
export function previewParserProfile(
sample_post: string,
profile: HeuristicProfile | Record<string, unknown>,
): Promise<{ fields: Record<string, string> }> {
return request<{ fields: Record<string, string> }>(
"/parser-builder/preview",
{ method: "POST", body: JSON.stringify({ sample_post, profile }) },
ADMIN_BASE,
);
}
export function createProfileJob(payload: {
channel: string;
profile: HeuristicProfile | Record<string, unknown>;
sample_post?: string;
limit?: number;
interval_seconds?: number;
is_active?: boolean;
}): Promise<ParseJob> {
return request<ParseJob>(
"/parser-builder/jobs",
{ method: "POST", body: JSON.stringify(payload) },
ADMIN_BASE,
);
}
@@ -7,6 +7,7 @@ const route = useRoute();
const navItems = [
{ to: "/", label: "Карта", exact: true },
{ to: "/parsers", label: "Парсеры" },
{ to: "/parser-builder", label: "Конструктор" },
{ to: "/events", label: "События" },
{ to: "/analytics", label: "Аналитика" },
{ to: "/consumers", label: "ПИ" },
@@ -5,6 +5,7 @@ import AnalyticsView from "../views/AnalyticsView.vue";
import ConsumersView from "../views/ConsumersView.vue";
import EventsView from "../views/EventsView.vue";
import MapViewPage from "../views/MapViewPage.vue";
import ParserBuilderView from "../views/ParserBuilderView.vue";
import ParsersView from "../views/ParsersView.vue";
const router = createRouter({
@@ -16,6 +17,7 @@ const router = createRouter({
children: [
{ path: "", name: "map", component: MapViewPage },
{ path: "parsers", name: "parsers", component: ParsersView },
{ path: "parser-builder", name: "parser-builder", component: ParserBuilderView },
{ path: "events", name: "events", component: EventsView },
{ path: "analytics", name: "analytics", component: AnalyticsView },
{ path: "consumers", name: "consumers", component: ConsumersView },
@@ -0,0 +1,320 @@
<script setup lang="ts">
import { onMounted, ref } from "vue";
import { useRouter } from "vue-router";
import {
createProfileJob,
fetchTargetFields,
generateParserProfile,
previewParserProfile,
type HeuristicProfile,
type TargetField,
} from "../api/parserBuilder";
const INTERVAL_OPTIONS = [
{ value: 900, label: "15 мин" },
{ value: 1800, label: "30 мин" },
{ value: 3600, label: "1 ч" },
{ value: 10800, label: "3 ч" },
{ value: 21600, label: "6 ч" },
{ value: 86400, label: "24 ч" },
];
const router = useRouter();
const samplePost = ref("");
const targetFields = ref<TargetField[]>([]);
const profile = ref<HeuristicProfile | null>(null);
const profileJson = ref("");
const previewFields = ref<Record<string, string> | null>(null);
const channel = ref("");
const limit = ref(100);
const intervalSeconds = ref(3600);
const loadingFields = ref(true);
const generating = ref(false);
const previewing = ref(false);
const saving = ref(false);
const error = ref("");
const success = ref("");
async function loadTargetFields() {
loadingFields.value = true;
try {
const data = await fetchTargetFields();
targetFields.value = data.fields;
error.value = "";
} catch (err) {
error.value = err instanceof Error ? err.message : "Не удалось загрузить целевые поля";
} finally {
loadingFields.value = false;
}
}
function syncProfileFromJson(): HeuristicProfile | null {
if (!profileJson.value.trim()) return null;
try {
const parsed = JSON.parse(profileJson.value) as HeuristicProfile;
profile.value = parsed;
error.value = "";
return parsed;
} catch {
error.value = "Некорректный JSON профиля";
return null;
}
}
async function handleGenerate() {
if (!samplePost.value.trim()) {
error.value = "Вставьте образец поста";
return;
}
generating.value = true;
error.value = "";
success.value = "";
previewFields.value = null;
try {
const data = await generateParserProfile(samplePost.value);
profile.value = data.profile;
profileJson.value = JSON.stringify(data.profile, null, 2);
success.value = "Профиль сгенерирован. Preview и runtime идут только по правилам — без LLM.";
} catch (err) {
error.value = err instanceof Error ? err.message : "Ошибка генерации";
} finally {
generating.value = false;
}
}
async function handlePreview() {
const current = syncProfileFromJson();
if (!current) {
if (!error.value) error.value = "Сначала сгенерируйте профиль";
return;
}
if (!samplePost.value.trim()) {
error.value = "Нужен образец поста для preview";
return;
}
previewing.value = true;
error.value = "";
try {
const data = await previewParserProfile(samplePost.value, current);
previewFields.value = data.fields;
} catch (err) {
error.value = err instanceof Error ? err.message : "Ошибка preview";
} finally {
previewing.value = false;
}
}
async function handleCreate() {
const current = syncProfileFromJson();
if (!current) {
if (!error.value) error.value = "Нужен профиль";
return;
}
if (!channel.value.trim()) {
error.value = "Укажите канал Telegram";
return;
}
saving.value = true;
error.value = "";
success.value = "";
try {
const job = await createProfileJob({
channel: channel.value.trim(),
profile: current,
sample_post: samplePost.value.trim() || undefined,
limit: limit.value,
interval_seconds: intervalSeconds.value,
is_active: true,
});
success.value = `Парсер #${job.id} создан (extract_mode=profile).`;
await router.push({ name: "parsers" });
} catch (err) {
error.value = err instanceof Error ? err.message : "Не удалось создать парсер";
} finally {
saving.value = false;
}
}
onMounted(loadTargetFields);
</script>
<template>
<div class="page">
<h2 class="page-heading">Конструктор парсера</h2>
<p class="intro">
Вставьте образец структурированного поста → DeepSeek один раз построит статичный профиль правил.
Дальше preview и парсинг канала идут только по профилю, без LLM на каждый пост.
</p>
<section class="card">
<h3>Образец поста</h3>
<textarea
v-model="samplePost"
class="sample-area"
rows="12"
placeholder="15.03.2024 Населённый пункт&#10;&#10;Описание события…&#10;&#10;48.123456, 37.654321&#10;#тема"
/>
<div class="form-actions">
<button
type="button"
class="btn btn-primary"
:disabled="generating || !samplePost.trim()"
@click="handleGenerate"
>
{{ generating ? "Генерация..." : "Сгенерировать парсер" }}
</button>
<button
type="button"
class="btn"
:disabled="previewing || !profileJson.trim()"
@click="handlePreview"
>
{{ previewing ? "Preview..." : "Preview" }}
</button>
</div>
</section>
<div class="grid-2">
<section class="card">
<h3>Целевая таблица <span class="muted">(Event / ingest)</span></h3>
<p v-if="loadingFields" class="muted">Загрузка…</p>
<div v-else class="table-wrap">
<table class="admin-table">
<thead>
<tr>
<th>Поле</th>
<th>Тип</th>
<th>Описание</th>
</tr>
</thead>
<tbody>
<tr v-for="field in targetFields" :key="field.name">
<td><code>{{ field.name }}</code></td>
<td>{{ field.type }}</td>
<td>{{ field.description }}</td>
</tr>
</tbody>
</table>
</div>
</section>
<section class="card">
<h3>Профиль (JSON)</h3>
<p v-if="!profileJson" class="muted">Появится после генерации. Можно слегка поправить правила.</p>
<textarea
v-model="profileJson"
class="profile-area"
rows="16"
:disabled="!profileJson && !profile"
spellcheck="false"
/>
</section>
</div>
<section v-if="previewFields" class="card">
<h3>Preview строки</h3>
<div class="table-wrap">
<table class="admin-table">
<thead>
<tr>
<th v-for="key in Object.keys(previewFields)" :key="key">{{ key }}</th>
</tr>
</thead>
<tbody>
<tr>
<td v-for="(val, key) in previewFields" :key="key" class="preview-cell">
{{ val || "—" }}
</td>
</tr>
</tbody>
</table>
</div>
</section>
<section class="card">
<h3>Создать Telegram-парсер</h3>
<form class="admin-form" @submit.prevent="handleCreate">
<div class="form-row">
<label>
Канал
<input v-model="channel" type="text" placeholder="example_channel" required />
</label>
<label>
Лимит batch
<input v-model.number="limit" type="number" min="1" max="1000" />
</label>
<label>
Интервал
<select v-model.number="intervalSeconds">
<option v-for="opt in INTERVAL_OPTIONS" :key="opt.value" :value="opt.value">
{{ opt.label }}
</option>
</select>
</label>
</div>
<div class="form-actions">
<button
type="submit"
class="btn btn-primary"
:disabled="saving || !profileJson.trim() || !channel.trim()"
>
{{ saving ? "Создание..." : "Создать парсер" }}
</button>
</div>
</form>
</section>
<p v-if="error" class="form-error">{{ error }}</p>
<p v-if="success" class="form-success">{{ success }}</p>
</div>
</template>
<style scoped>
.intro {
margin: 0 0 1rem;
font-size: 0.9rem;
color: #555;
line-height: 1.5;
max-width: 72ch;
}
.sample-area,
.profile-area {
width: 100%;
font: inherit;
font-family: ui-monospace, SFMono-Regular, Menlo, Consolas, monospace;
font-size: 0.8rem;
padding: 0.6rem 0.75rem;
border: 1px solid #d1d5db;
border-radius: 6px;
resize: vertical;
line-height: 1.45;
}
.grid-2 {
display: grid;
grid-template-columns: 1fr 1fr;
gap: 1rem;
}
@media (max-width: 960px) {
.grid-2 {
grid-template-columns: 1fr;
}
}
.preview-cell {
max-width: 180px;
white-space: pre-wrap;
word-break: break-word;
font-size: 0.8rem;
}
.form-success {
color: #2e7d32;
font-size: 0.875rem;
margin-top: 0.5rem;
}
</style>
@@ -79,9 +79,15 @@ function parseLines(text: string): string[] {
.filter(Boolean);
}
function extractModeLabel(cfg: Record<string, unknown>): string {
if (cfg.extract_mode === "llm") return " [LLM]";
if (cfg.extract_mode === "profile") return " [profile]";
return "";
}
function sourceSummary(job: ParseJob): string {
const cfg = job.source_config;
const mode = cfg.extract_mode === "llm" ? " [LLM]" : "";
const mode = extractModeLabel(cfg);
if (job.source_type === "telegram") {
return (channelFromConfig(cfg) || "—") + mode;
}
@@ -92,6 +98,21 @@ function sourceSummary(job: ParseJob): string {
return "—";
}
function isProfileJob(job: ParseJob | null): boolean {
return !!job && job.source_config.extract_mode === "profile";
}
function profileJson(job: ParseJob | null): string {
if (!job) return "";
const profile = job.source_config.heuristic_profile;
if (!profile || typeof profile !== "object") return "";
try {
return JSON.stringify(profile, null, 2);
} catch {
return "";
}
}
function formatInterval(seconds: number): string {
const opt = INTERVAL_OPTIONS.find((item) => item.value === seconds);
if (opt) return opt.label;
@@ -111,8 +132,19 @@ function buildSourceConfig(
input_mode: "urls" | "texts" | "mixed";
extract_mode: "heuristic" | "llm";
},
existing?: Record<string, unknown>,
): Record<string, unknown> {
if (sourceType === "telegram") {
// Preserve builder-created profile jobs (edit form has no profile editor)
if (existing?.extract_mode === "profile" && existing.heuristic_profile) {
return {
channel: data.channel.trim(),
limit: data.limit,
extract_mode: "profile",
heuristic_profile: existing.heuristic_profile,
sample_post: existing.sample_post ?? null,
};
}
return {
channel: data.channel.trim(),
limit: data.limit,
@@ -242,7 +274,7 @@ async function handleSaveEdit() {
error.value = "";
try {
await updateJob(editingJob.value.id, {
source_config: buildSourceConfig(sourceType, editForm.value),
source_config: buildSourceConfig(sourceType, editForm.value, editingJob.value.source_config),
interval_seconds: editForm.value.interval_seconds,
is_active: editForm.value.is_active,
});
@@ -318,6 +350,8 @@ onUnmounted(() => {
<p class="hint">
{{ formHint }}
Дубликаты по <code>source_url</code> не записываются.
Статичный парсер по образцу поста —
<router-link to="/parser-builder">Конструктор</router-link>.
</p>
<form class="admin-form" @submit.prevent="handleSubmit">
<div class="form-row">
@@ -479,13 +513,21 @@ onUnmounted(() => {
Лимит
<input v-model.number="editForm.limit" type="number" min="1" max="1000" />
</label>
<label>
<label v-if="!isProfileJob(editingJob)">
Извлечение
<select v-model="editForm.extract_mode">
<option value="heuristic">Эвристики</option>
<option value="llm">LLM DeepSeek</option>
</select>
</label>
<p v-else class="hint profile-hint">
Режим: <code>profile</code> (статичный профиль из конструктора).
Сменить правила — через повторную генерацию в «Конструктор».
</p>
<label v-if="isProfileJob(editingJob)" class="span-2">
Профиль (только просмотр)
<textarea class="profile-ro" :value="profileJson(editingJob)" rows="8" readonly />
</label>
</template>
<template v-else-if="editingJob.source_type === 'crawl4ai'">
<label class="span-2">
@@ -611,4 +653,17 @@ onUnmounted(() => {
gap: 0.5rem;
margin-top: 1.5rem;
}
.profile-hint {
flex: 1 1 100%;
margin: 0;
}
.profile-ro {
width: 100%;
font-family: ui-monospace, SFMono-Regular, Menlo, Consolas, monospace;
font-size: 0.75rem;
background: #f8f8f8;
color: #444;
}
</style>
+9
View File
@@ -63,6 +63,15 @@ async def run(job_id, source_config, *, ctx) -> tuple[list[dict], str | None]
- Crawl4AI: страница → `LLMExtractionStrategy` (DeepSeek) с fallback на тот же DeepSeek по markdown;
- в UI «Парсеры»: поле **Извлечение** = LLM DeepSeek.
### Profile-режим (`extract_mode: profile`)
Статичный парсер из конструктора ЦА:
- `heuristic_profile` в `source_config` (схема `contracts/heuristic_profile.py`);
- интерпретатор: `workers/heuristic_profile.py` (те же правила, что preview в ЦА);
- Telegram batch + listener применяют профиль без DeepSeek;
- генерация профиля — только в админке (`/admin/parser-builder/generate`).
```mermaid
flowchart LR
Job[JobPayload]
@@ -7,6 +7,7 @@ import logging
from contracts.sources import TelegramSourceConfig
from workers.adapters.base import WorkerContext
from workers.converter import event_record_to_ingest
from workers.heuristic_profile import extract_with_profile
from workers.llm_extract import (
DEFAULT_EXTRACT_SCHEMA,
DEFAULT_INSTRUCTION,
@@ -57,11 +58,37 @@ class TelegramAdapter:
return [], "extract_mode=llm requires DEEPSEEK_API_KEY in worker env"
return await _extract_posts_with_llm(posts, cfg)
if cfg.extract_mode == "profile":
return _extract_posts_with_profile(posts, cfg)
records = parse_event_posts(posts)
events = [event_record_to_ingest(r) for r in records]
return events, None
def _extract_posts_with_profile(posts, cfg: TelegramSourceConfig) -> tuple[list[dict], str | None]:
profile = cfg.heuristic_profile or {}
events: list[dict] = []
for post in posts:
text = (post.text or "").strip()
if not text:
continue
events.append(
extract_with_profile(
text,
profile,
source_url=post.url,
source_type="telegram",
extra_metadata={
"channel": post.channel,
"message_id": post.id,
"post_date": post.date.isoformat() if post.date else None,
},
)
)
return events, None
async def _extract_posts_with_llm(posts, cfg: TelegramSourceConfig) -> tuple[list[dict], str | None]:
schema = cfg.extract_schema or DEFAULT_EXTRACT_SCHEMA
instruction = cfg.instruction or DEFAULT_INSTRUCTION
@@ -0,0 +1,73 @@
"""Apply HeuristicProfile to Telegram posts → IngestEventItem-shaped dicts."""
from __future__ import annotations
from typing import Any
from contracts.heuristic_profile import HeuristicProfile, apply_profile
from workers.llm_extract import parse_coords, parse_date
def profile_fields_to_ingest(
*,
source_type: str,
source_url: str,
raw_text: str,
fields: dict[str, str],
extra_metadata: dict | None = None,
) -> dict:
lat, lng = parse_coords(str(fields.get("coords") or ""))
description = str(fields.get("description") or raw_text)[:8000]
locality = str(fields.get("locality") or "")
region = str(fields.get("region") or locality or "") or None
if fields.get("title"):
title = str(fields["title"])
elif locality:
title = locality
elif description:
title = description.splitlines()[0][:120]
else:
title = ""
topic = str(fields.get("topic") or "telegram_profile")
event_date = parse_date(str(fields.get("event_date") or ""))
meta: dict[str, Any] = {
"extract_mode": "profile",
"extracted": dict(fields),
}
if extra_metadata:
meta.update(extra_metadata)
return {
"source_type": source_type,
"source_url": source_url,
"raw_text": raw_text[:20000],
"title": title[:255],
"description": description,
"locality": locality,
"latitude": lat,
"longitude": lng,
"event_date": event_date.isoformat() if event_date else None,
"region": region,
"topic": topic,
"tags": [source_type, "profile"],
"metadata": meta,
}
def extract_with_profile(
text: str,
profile: HeuristicProfile | dict[str, Any],
*,
source_url: str,
source_type: str = "telegram",
extra_metadata: dict | None = None,
) -> dict:
fields = apply_profile(text, profile)
return profile_fields_to_ingest(
source_type=source_type,
source_url=source_url,
raw_text=text,
fields=fields,
extra_metadata=extra_metadata,
)
@@ -9,6 +9,7 @@ import httpx
from telethon import TelegramClient, events
from workers.converter import event_record_to_ingest
from workers.heuristic_profile import extract_with_profile
from workers.parsers.telegram_events import parse_event_post
from workers.sources.telegram_client import message_to_post, normalize_channel
@@ -22,11 +23,12 @@ REFRESH_SECONDS = int(os.getenv("TELEGRAM_LISTENER_REFRESH_SECONDS", "60"))
class TelegramListener:
def __init__(self, client: TelegramClient) -> None:
self.client = client
self._channels: dict[str, int] = {}
# channel -> {job_id, source_config}
self._channels: dict[str, dict[str, Any]] = {}
self._chat_ids: set[int] = set()
self._handlers_registered = False
async def fetch_subscriptions(self) -> dict[str, int]:
async def fetch_subscriptions(self) -> dict[str, dict[str, Any]]:
async with httpx.AsyncClient(timeout=30.0) as http:
response = await http.get(
f"{CA_API_URL}/internal/listener/subscriptions",
@@ -35,7 +37,7 @@ class TelegramListener:
response.raise_for_status()
data = response.json()
channels: dict[str, int] = {}
channels: dict[str, dict[str, Any]] = {}
for item in data:
raw = item.get("channel")
job_id = item.get("job_id")
@@ -46,7 +48,13 @@ class TelegramListener:
except ValueError:
logger.warning("Skip invalid channel in subscription: %r", raw)
continue
channels.setdefault(key, int(job_id))
channels.setdefault(
key,
{
"job_id": int(job_id),
"source_config": dict(item.get("source_config") or {}),
},
)
return channels
async def refresh_subscriptions(self) -> None:
@@ -73,10 +81,34 @@ class TelegramListener:
", ".join(sorted(channels)) or "(none)",
)
async def _ingest_post(self, channel: str, post) -> None:
def _build_event(self, channel: str, post) -> dict:
sub = self._channels.get(normalize_channel(channel)) or {}
cfg = sub.get("source_config") or {}
extract_mode = cfg.get("extract_mode") or "heuristic"
text = (post.text or "").strip()
if extract_mode == "profile" and cfg.get("heuristic_profile"):
return extract_with_profile(
text,
cfg["heuristic_profile"],
source_url=post.url,
source_type="telegram",
extra_metadata={
"channel": post.channel,
"message_id": post.id,
"post_date": post.date.isoformat() if post.date else None,
"listener": True,
},
)
# llm jobs fall back to heuristic in listener (LLM is batch-only by design)
record = parse_event_post(post)
event = event_record_to_ingest(record)
job_id = self._channels.get(normalize_channel(channel))
return event_record_to_ingest(record)
async def _ingest_post(self, channel: str, post) -> None:
event = self._build_event(channel, post)
sub = self._channels.get(normalize_channel(channel)) or {}
job_id = sub.get("job_id")
async with httpx.AsyncClient(timeout=60.0) as http:
response = await http.post(
+165
View File
@@ -0,0 +1,165 @@
"""Declarative heuristic profile: schema + static rule interpreter (CA preview + CP runtime)."""
from __future__ import annotations
import re
from typing import Any, Literal
from pydantic import BaseModel, Field, field_validator, model_validator
TARGET_FIELDS: tuple[str, ...] = (
"title",
"description",
"locality",
"event_date",
"coords",
"topic",
"region",
)
Strategy = Literal["regex", "line", "after_marker", "between", "full_text", "literal"]
class FieldRule(BaseModel):
strategy: Strategy
# regex: named or group(1); line: 0-based index; after_marker/between: markers
pattern: str | None = None
group: int = 1
line_index: int | None = None
marker: str | None = None
end_marker: str | None = None
value: str | None = None # literal
flags: str = "" # e.g. "im" → re.I|re.M
strip: bool = True
@field_validator("pattern", "marker", "end_marker", "value", mode="before")
@classmethod
def empty_to_none(cls, value: Any) -> Any:
if value is None:
return None
if isinstance(value, str) and not value.strip():
return None
return value
class HeuristicProfile(BaseModel):
version: Literal[1] = 1
fields: dict[str, FieldRule] = Field(default_factory=dict)
notes: str = ""
@model_validator(mode="after")
def known_fields_only(self) -> "HeuristicProfile":
unknown = set(self.fields) - set(TARGET_FIELDS)
if unknown:
raise ValueError(f"Unknown profile fields: {sorted(unknown)}")
return self
def target_field_specs() -> list[dict[str, str]]:
"""Fixed target table for parser-builder UI (roadmap: custom tables later)."""
return [
{"name": "title", "type": "string", "description": "Short event title"},
{"name": "description", "type": "string", "description": "Event summary / body"},
{"name": "locality", "type": "string", "description": "Place / settlement name"},
{
"name": "event_date",
"type": "string",
"description": "Date as DD.MM.YYYY or YYYY-MM-DD",
},
{
"name": "coords",
"type": "string",
"description": "Latitude, longitude if present",
},
{"name": "topic", "type": "string", "description": "Short topic tag"},
{"name": "region", "type": "string", "description": "Region (optional)"},
]
def _compile_flags(flags: str) -> int:
mapping = {
"i": re.IGNORECASE,
"m": re.MULTILINE,
"s": re.DOTALL,
}
result = 0
for ch in (flags or "").lower():
result |= mapping.get(ch, 0)
return result
def _apply_rule(text: str, rule: FieldRule) -> str:
raw = text or ""
value = ""
if rule.strategy == "literal":
value = rule.value or ""
elif rule.strategy == "full_text":
value = raw
elif rule.strategy == "line":
lines = raw.splitlines()
idx = 0 if rule.line_index is None else rule.line_index
if 0 <= idx < len(lines):
value = lines[idx]
elif rule.strategy == "regex":
if not rule.pattern:
return ""
match = re.search(rule.pattern, raw, _compile_flags(rule.flags))
if match:
try:
value = match.group(rule.group)
except IndexError:
value = match.group(0)
elif rule.strategy == "after_marker":
marker = rule.marker or ""
if not marker:
return ""
pos = raw.find(marker)
if pos < 0:
return ""
start = pos + len(marker)
rest = raw[start:]
if rule.end_marker:
end = rest.find(rule.end_marker)
value = rest[:end] if end >= 0 else rest
elif rule.pattern:
match = re.search(rule.pattern, rest, _compile_flags(rule.flags))
if match:
try:
value = match.group(rule.group)
except IndexError:
value = match.group(0)
else:
# first non-empty line after marker
for line in rest.splitlines():
if line.strip():
value = line
break
elif rule.strategy == "between":
start_m = rule.marker or ""
end_m = rule.end_marker or ""
if not start_m or not end_m:
return ""
start = raw.find(start_m)
if start < 0:
return ""
start += len(start_m)
end = raw.find(end_m, start)
if end < 0:
return ""
value = raw[start:end]
if rule.strip:
value = value.strip()
return value
def apply_profile(text: str, profile: HeuristicProfile | dict[str, Any]) -> dict[str, str]:
"""Apply static rules to post text → string field map (no LLM)."""
if isinstance(profile, dict):
profile = HeuristicProfile.model_validate(profile)
result: dict[str, str] = {name: "" for name in TARGET_FIELDS}
for name, rule in profile.fields.items():
result[name] = _apply_rule(text, rule)
return result
+15 -2
View File
@@ -10,16 +10,29 @@ from pydantic import BaseModel, Field, field_validator, model_validator
class TelegramSourceConfig(BaseModel):
channel: str = Field(min_length=1)
limit: int = Field(default=100, ge=1, le=1000)
# heuristic = legacy telegram_events parser; llm = DeepSeek structured extract
extract_mode: Literal["heuristic", "llm"] = "heuristic"
# heuristic = legacy telegram_events; llm = DeepSeek per post; profile = static rules
extract_mode: Literal["heuristic", "llm", "profile"] = "heuristic"
extract_schema: dict[str, str] | None = None
instruction: str | None = None
heuristic_profile: dict | None = None
sample_post: str | None = None # audit / re-generate; not required at runtime
@field_validator("channel")
@classmethod
def strip_channel(cls, value: str) -> str:
return value.strip()
@model_validator(mode="after")
def profile_requires_rules(self) -> "TelegramSourceConfig":
if self.extract_mode != "profile":
return self
if not self.heuristic_profile:
raise ValueError("heuristic_profile required when extract_mode=profile")
from contracts.heuristic_profile import HeuristicProfile
HeuristicProfile.model_validate(self.heuristic_profile)
return self
class Crawl4AISourceConfig(BaseModel):
urls: list[str] = Field(min_length=1)
+2
View File
@@ -22,6 +22,8 @@ services:
build:
context: .
dockerfile: centers/analytics/api/Dockerfile
env_file:
- .env
environment:
DATABASE_URL: postgresql://mapmil:mapmil@ca-db:5432/mapmil
REDIS_URL: redis://redis:6379/0
+4 -1
View File
@@ -184,7 +184,9 @@ async def run(job_id, source_config, *, ctx) -> tuple[list[dict], str | None]
Опционально: `extract_mode: llm` (DeepSeek) для telegram/crawl4ai в batch.
Подробности: [centers/parsing/ARCHITECTURE.md](../centers/parsing/ARCHITECTURE.md).
**Конструктор парсера** (`/parser-builder`): менеджер вставляет образец поста → DeepSeek один раз строит `HeuristicProfile` (статичные regex/правила) → job с `extract_mode=profile`. Runtime (batch + listener) применяет только профиль, без LLM. Кастомные пользовательские таблицы — roadmap (подмена `target-fields`).
Подробности: [centers/parsing/ARCHITECTURE.md](../centers/parsing/ARCHITECTURE.md), [centers/analytics/ARCHITECTURE.md](../centers/analytics/ARCHITECTURE.md).
### 2.4. Distribution (ПИ)
@@ -206,6 +208,7 @@ async def run(job_id, source_config, *, ctx) -> tuple[list[dict], str | None]
| `contracts/ingest.py` | `IngestEventItem` / payload ingest |
| `contracts/sources.py` | Валидация `source_config` по `source_type` |
| `contracts/queues.py` | `SOURCE_FAMILY` → ключ очереди |
| `contracts/heuristic_profile.py` | Статичный профиль конструктора + `apply_profile` |
Правило: меняете форму события / конфиг источника / очередь — сначала `contracts/`, потом CA/CP/UI.
+3 -1
View File
@@ -57,12 +57,14 @@ Listener добавляет флаг `listener: true` на стороне CA API
| source_type | Модель | Главные поля |
|-------------|--------|--------------|
| `telegram` | `TelegramSourceConfig` | `channel`, `limit`, `extract_mode` (`heuristic`\|`llm`) |
| `telegram` | `TelegramSourceConfig` | `channel`, `limit`, `extract_mode` (`heuristic`\|`llm`\|`profile`), `heuristic_profile` |
| `crawl4ai` | `Crawl4AISourceConfig` | `urls`, `extract_mode`, `extract_schema`, `domain_profile` |
| `viina` | `ViinaSourceConfig` | `urls` / `texts`, `input_mode` |
Реестр: `CONFIG_MODELS` + `parse_source_config(source_type, raw)`.
Профиль конструктора: `contracts/heuristic_profile.py` (`HeuristicProfile`, `apply_profile`).
---
## Очереди (`queues.py`)
+2 -1
View File
@@ -5,7 +5,7 @@
- Docker + Docker Compose
- Файл `.env` (из `.env.example`)
- Для Telegram: `data/telegram.session` + `TELEGRAM_API_ID` / `TELEGRAM_API_HASH`
- Опционально: `DEEPSEEK_API_KEY` для `extract_mode: llm`
- Опционально: `DEEPSEEK_API_KEY` для `extract_mode: llm` (воркеры) и конструктора парсера на `ca-api`
## Быстрый старт
@@ -109,6 +109,7 @@ docker compose logs -f cp-workers-web
| Frontend 000 / нет контейнеров | `docker compose up -d` (без `--build`, если registry недоступен, но образы уже есть) |
| Job не берётся | Смотреть `ENABLED_ADAPTERS` / `WORKER_FAMILIES` нужного сервиса |
| LLM не работает | `DEEPSEEK_API_KEY` в `.env`, перезапуск `cp-workers` / `cp-workers-web` |
| Generate в конструкторе 503 | `DEEPSEEK_API_KEY` в `.env` + `env_file` у `ca-api`, перезапуск `ca-api` |
| Изменения UI не видны | Пересобрать `ca-frontend` |
## Границы при разработке