Compare commits

...
9 Commits
Author SHA1 Message Date
gitrusprusandCursor 3f9dc6643b Add reusable LLM parser profiles with multi-event extract.
Support kind=llm profiles (instruction/schema), optional multi-event posts via #eN URLs, and recover stale running/queued parse jobs after worker crashes.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-13 20:22:38 +03:00
gitrusprusandCursor 5811ecb134 Add VPN admin tab with subscription proxy via cp-vpn for selected sources.
Persist settings in CA, expose /admin/vpn and /internal/vpn, and route Telegram (and optional web/nlp) traffic through mihomo SOCKS when enabled.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-16 23:41:33 +03:00
gitrusprusandCursor 8ab606747f Add public map with JWT admin login and protect admin APIs.
Keep map reads open; gate admin UI/nav and object mutations behind env-based admin credentials, and default parser batch limit to 10.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-16 23:04:21 +03:00
gitrusprusandCursor 52185aeaa8 Add profile generate refine via manager hint and auto-preview.
When Preview leaves fields empty, managers can send a hint plus the current profile so DeepSeek revises rules; generate now returns preview and empty_fields for the UI.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-16 22:25:27 +03:00
gitrusprusandCursor 71f8cf169e Add Profile/Channel entities with CRUD and remove legacy parser builder.
Introduce reusable ParserProfile and ParseChannel, pair jobs with enqueue flatten, Events admin CRUD, and drop inline/legacy parser-builder UI and aliases.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-16 14:51:03 +03:00
gitrusprus e196de7320 Merge branch 'Генерация_парсера(пост_+_БД)'
Parser builder with extract_mode=profile for Telegram.
2026-08-16 14:14:25 +03:00
gitrusprusandCursor 699e9be503 Add parser builder: one-shot DeepSeek profile for Telegram extract_mode=profile.
Generate static HeuristicProfile in CA admin, preview and run without LLM on each post via shared interpreter in CP batch and listener.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-16 14:14:18 +03:00
gitrusprusandCursor 6dff3c1c3d Document platform architecture and Telegram proxy local setup.
Add developer docs for CA/CP/PI flows and wire cp-workers to the host SOCKS proxy so local Telegram auth and parsing work reliably.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-15 22:57:07 +03:00
gitrusprusandCursor 5f8ed25759 Add force-run button for parsers in the admin UI.
Allow re-enqueueing any idle parse job so operators can trigger a post fetch without waiting for the schedule.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-15 22:35:35 +03:00
67 changed files with 6759 additions and 537 deletions
+15 -2
View File
@@ -3,20 +3,33 @@ TELEGRAM_API_ID=12345678
TELEGRAM_API_HASH=your_api_hash_here
TELEGRAM_SESSION_PATH=/data/telegram.session
# Optional Telegram proxy
# Optional Telegram proxy fallback (when VPN tab is off)
# Prefer Admin UI → VPN (subscription / socks5 / http).
# TELEGRAM_PROXY_TYPE=socks5
# TELEGRAM_PROXY_HOST=127.0.0.1
# TELEGRAM_PROXY_PORT=1080
# TELEGRAM_PROXY_USER=
# TELEGRAM_PROXY_PASS=
# Platform internals
INTERNAL_TOKEN=dev-internal-token
TEST_PI_API_KEY=test-pi-api-key-change-me
# DeepSeek LLM for extract_mode=llm (Telegram unstructured + Crawl4AI LLM)
# Admin UI login (public map stays open; /admin and object writes need JWT)
ADMIN_USER=admin
ADMIN_PASSWORD=change-me
ADMIN_JWT_SECRET=change-me-jwt-secret
# DeepSeek LLM:
# - extract_mode=llm on CP workers (Telegram unstructured + Crawl4AI)
# - Generate профиля парсера на ca-api (один раз; runtime — только правила)
# DEEPSEEK_API_KEY=sk-...
# DEEPSEEK_BASE_URL=https://api.deepseek.com
# DEEPSEEK_MODEL=deepseek-chat
# Recover parse jobs stuck in running/queued after worker crash (seconds, default 900)
# STALE_JOB_SECONDS=900
# CP adapter workers (set in docker-compose; override locally if needed)
# ENABLED_ADAPTERS=telegram
# WORKER_FAMILIES=telegram
+13 -2
View File
@@ -36,10 +36,21 @@ docker compose up --build cp-workers-web # single worker rebuild
Health: `curl http://localhost:8080/api/health`
Admin auth (map is public; `/admin/*` needs JWT from `ADMIN_USER` / `ADMIN_PASSWORD`):
```bash
TOKEN=$(curl -s -X POST http://localhost:8080/admin/auth/login \
-H 'Content-Type: application/json' \
-d '{"username":"admin","password":"change-me"}' | jq -r .access_token)
```
VPN (subscription / SOCKS): Admin UI → **VPN**, or `PUT /admin/vpn`. Workers and `cp-vpn` read `GET /internal/vpn`.
Create Telegram parser:
```bash
curl -X POST http://localhost:8080/admin/jobs \
-H "Authorization: Bearer $TOKEN" \
-H 'Content-Type: application/json' \
-d '{"source_type":"telegram","source_config":{"channel":"example","limit":50}}'
```
@@ -63,7 +74,7 @@ Adapters return dicts matching `IngestEventItem` (`contracts/ingest.py`). Requir
### LLM extract (runtime, not dev)
`extract_mode: llm` uses DeepSeek via `workers/llm_extract.py`. Key: `DEEPSEEK_API_KEY` in `.env`. Do not confuse with Cursor dev agents.
`extract_mode: llm` uses DeepSeek via `workers/llm_extract.py` (batch). Reusable profiles in admin UI (`kind=heuristic|llm`) flatten into `extract_mode=profile` or `llm`. Key: `DEEPSEEK_API_KEY` in `.env`. Do not confuse with Cursor dev agents.
## Common tasks
@@ -86,7 +97,7 @@ Adapters return dicts matching `IngestEventItem` (`contracts/ingest.py`). Requir
- Rule: `.cursor/rules/platform.mdc` (always on)
- Skill: `.cursor/skills/add-parser-adapter/` for end-to-end new sources
- Deep docs: `README.md`, `centers/parsing/ARCHITECTURE.md`
- Deep docs: `docs/` (overview + data-flow), `centers/analytics/ARCHITECTURE.md`, `centers/parsing/ARCHITECTURE.md`
## Commit style
+35 -16
View File
@@ -2,6 +2,8 @@
Единая платформа: ЦА (аналитика и карта) + ЦП (адаптеры парсинга: Telegram, Crawl4AI, VIINA).
**Документация для разработчиков:** [`docs/`](docs/README.md) — [обзор архитектуры](docs/architecture-overview.md), [поток данных](docs/data-flow.md), [локальный запуск](docs/local-dev.md).
## Архитектура
```mermaid
@@ -15,13 +17,14 @@ flowchart LR
CA -->|GET /api/v1/events| PI
```
| Центр | Контейнеры | Назначение |
|-------|------------|------------|
| **ЦА** | `ca-db`, `ca-api`, `ca-frontend` | PostgreSQL, ingest API, карта, distribution API |
| **ЦП** | `cp-workers`, `cp-workers-web`, `cp-workers-nlp` | Адаптеры: Telegram / Crawl4AI / VIINA |
| **Общее** | `redis` | Очереди `cp:jobs:{telegram\|web\|nlp}` |
| Центр | Контейнеры | Назначение | Документ |
|-------|------------|------------|----------|
| **ЦА** | `ca-db`, `ca-api`, `ca-frontend` | PostgreSQL, ingest API, карта, distribution API | [`centers/analytics/ARCHITECTURE.md`](centers/analytics/ARCHITECTURE.md) |
| **ЦП** | `cp-workers`, `cp-workers-web`, `cp-workers-nlp` | Адаптеры: Telegram / Crawl4AI / VIINA | [`centers/parsing/ARCHITECTURE.md`](centers/parsing/ARCHITECTURE.md) |
| **ПИ** | (маршрут в ЦА) | `GET /api/v1/events` | [`centers/analytics/DISTRIBUTION.md`](centers/analytics/DISTRIBUTION.md) |
| **Общее** | `redis` | Очереди `cp:jobs:{telegram\|web\|nlp}` | [`docs/contracts.md`](docs/contracts.md) |
Подробности ЦП: [`centers/parsing/ARCHITECTURE.md`](centers/parsing/ARCHITECTURE.md).
Подробный обзор платформы: [`docs/architecture-overview.md`](docs/architecture-overview.md).
## Структура monorepo
@@ -29,13 +32,16 @@ flowchart LR
MapMil/
├── centers/
│ ├── analytics/
│ │ ├── api/ # CA backend (FastAPI + PostgreSQL)
│ │ └── frontend/ # CA admin UI (Vue + Leaflet)
│ │ ├── api/ # CA backend (FastAPI + PostgreSQL)
│ │ ├── frontend/ # CA admin UI (Vue + Leaflet)
│ │ ├── ARCHITECTURE.md
│ │ └── DISTRIBUTION.md # ПИ
│ └── parsing/
│ ├── ARCHITECTURE.md
│ └── workers/ # CP workers + adapters
├── contracts/ # Shared schemas (ingest, jobs, sources, queues)
├── data/ # telegram.session (локально, не в git)
│ └── workers/ # CP workers + adapters
├── contracts/ # Shared schemas (ingest, jobs, sources, queues)
├── docs/ # Документация для разработчиков
├── data/ # telegram.session (локально, не в git)
├── docker-compose.yml
└── .env
```
@@ -61,20 +67,23 @@ cp ../SocialParser/data/telegram.session data/
docker compose up --build
```
4. Откройте UI: [http://localhost:8080](http://localhost:8080)
4. Откройте UI: [http://localhost:8080](http://localhost:8080) — карта публичная; разделы админки — после входа (`ADMIN_USER` / `ADMIN_PASSWORD` из `.env`).
## Admin UI
Веб-интерфейс ЦА доступен по тем же адресу. Навигация в шапке:
Веб-интерфейс ЦА доступен по тому же адресу. Карта (`/`) открыта всем. Остальные пункты навигации видны после логина (`/login`).
| Раздел | Путь | Описание |
|--------|------|----------|
| **Карта** | `/` | Интерактивная карта событий: навигация по датам (flatpickr), пресеты периода, фильтры региона/темы/источника, подложки Яндекс/OSM/Topo/ESRI, линейка, полноэкранный режим, центрирование по координатам и городам, поиск населённых пунктов (Nominatim). CRUD для ручных объектов (ПКМ). Поддерживает `?eventId=` |
| **Парсеры** | `/parsers` | Адаптеры `telegram` / `crawl4ai` / `viina`, интервал, CRUD; дедуп по `source_url` |
| **Карта** | `/` | Интерактивная карта событий (публичный просмотр). CRUD ручных объектов — только для админа (ПКМ). Поддерживает `?eventId=` |
| **Парсеры** | `/parsers` | Связка канал + профиль (`heuristic`→`extract_mode=profile`, `llm`→`extract_mode=llm`); дедуп по `source_url` |
| **Профили** | `/parser-profiles` | Reusable heuristic (правила) или llm (instruction/schema); целевые поля = Event |
| **События** | `/events` | Фильтрация, пагинация, просмотр деталей, ссылка «На карте» для событий с координатами |
| **Аналитика** | `/analytics` | KPI-карточки, график динамики ingest за 30 дней, топ населённых пунктов и регионов |
| **ПИ** | `/consumers` | CRUD подписчиков distribution API, ротация ключей, тест среза через `/api/v1/events` |
Авторизация админки: `POST /admin/auth/login` → JWT Bearer. Все `/admin/*` (кроме login) и мутации `/api/objects` требуют токен. `GET /api/map/*` публичен.
## Сохранение Telegram-сессии
**Важно:** существующий файл сессии **не удаляется и не пересоздаётся**.
@@ -126,10 +135,15 @@ docker compose up --build
| PATCH | `/admin/consumers/{id}` | Обновить подписчика |
| POST | `/admin/consumers/{id}/rotate-key` | Сменить API-ключ |
Пример задания Telegram:
Пример задания Telegram (нужен JWT после `POST /admin/auth/login`):
```bash
TOKEN=$(curl -s -X POST http://localhost:8080/admin/auth/login \
-H 'Content-Type: application/json' \
-d '{"username":"admin","password":"change-me"}' | jq -r .access_token)
curl -X POST http://localhost:8080/admin/jobs \
-H "Authorization: Bearer $TOKEN" \
-H 'Content-Type: application/json' \
-d '{"source_type":"telegram","source_config":{"channel":"creamy_caprice","limit":50}}'
```
@@ -201,3 +215,8 @@ docker compose down
```
Данные PostgreSQL сохраняются в volume `pgdata`.
## Вход в админку (/login):
Логин: admin
Пароль: change-me
+182
View File
@@ -0,0 +1,182 @@
# Архитектура ЦА (Analytics Center)
Центр аналитики — ядро MapMil: PostgreSQL, FastAPI, Vue admin/карта, distribution API для ПИ.
Общий контекст: [docs/architecture-overview.md](../../docs/architecture-overview.md).
ПИ отдельно: [DISTRIBUTION.md](DISTRIBUTION.md).
---
## Общими чертами
ЦА:
- хранит события, jobs, объекты карты, consumers;
- отдаёт единственный UI (`ca-frontend`);
- ставит задания в Redis для ЦП;
- принимает ingest от ЦП;
- отдаёт срезы внешним системам через `/api/v1/events`.
```mermaid
flowchart TB
UI[ca-frontend Vue]
API[ca-api FastAPI]
DB[(PostgreSQL)]
Redis[(Redis)]
CP[CP workers]
UI --> API
API --> DB
API -->|enqueue| Redis
Redis --> CP
CP -->|/internal/*| API
```
Контейнеры: `ca-db`, `ca-api`, `ca-frontend` (+ общий `redis`).
---
## Подробнее
### Структура
```text
centers/analytics/
├── ARCHITECTURE.md
├── DISTRIBUTION.md
├── api/
│ ├── Dockerfile
│ ├── requirements.txt
│ └── app/
│ ├── main.py # lifespan: migrations, scheduler, seed
│ ├── database.py
│ ├── models.py
│ ├── schemas.py
│ ├── deps.py # X-Internal-Token
│ ├── seed.py
│ ├── storage.py # uploads
│ ├── routers/
│ │ ├── objects.py # /api/health, /api/objects, media
│ │ ├── map.py # /api/map/*
│ │ ├── admin.py # /admin/* (jobs, events, consumers)
│ │ ├── parser_profiles.py # /admin/parser-profiles/*
│ │ ├── parse_channels.py # /admin/parse-channels/*
│ │ ├── internal.py # /internal/* (только ЦП)
│ │ └── v1.py # /api/v1/* (ПИ)
│ └── services/
│ ├── jobs.py # Redis RPUSH + flatten Profile/Channel
│ ├── scheduler.py # периодический re-queue
│ ├── ingest.py # дедуп + map sync
│ ├── parser_builder.py # DeepSeek → HeuristicProfile (один раз)
│ ├── filtering.py
│ ├── map_query.py
│ ├── events_query.py
│ ├── analytics.py
│ └── migrations.py
└── frontend/
├── Dockerfile # Vite build + nginx
├── nginx.conf # proxy /api /admin /internal → ca-api
└── src/
├── views/ # Map, Channels, Profiles, Parsers, …
├── components/ # карта, CRUD объектов
├── api/ # HTTP-клиенты
└── router/index.ts
```
### Модели данных
| Модель | Таблица | Назначение |
|--------|--------|------------|
| `Event` | `events` | Нормализованное событие; UK `source_url` |
| `ParserProfile` | `parser_profiles` | Статичный heuristic-профиль (JSON-правила) |
| `ParseChannel` | `parse_channels` | Канал (Telegram handle) |
| `ParseJob` | `parse_jobs` | Связка канал+профиль, интервал, статус; FK `profile_id` + `channel_id` |
| `MapObject` | `map_objects` | Точка на карте (event или ручная) |
| `ObjectMedia` | `object_media` | Файлы к объектам |
| `Consumer` | `consumers` | Подписчик ПИ (hash ключа) |
| `ConsumerFilter` | `consumer_filters` | Фильтры среза ПИ |
### HTTP-поверхности
| Prefix | Кто вызывает | Содержание |
|--------|--------------|------------|
| `/api/*` | UI, публичный health | Карта, объекты, медиа |
| `/admin/*` | UI admin | Jobs, channels, profiles, events, analytics, consumers |
| `/internal/*` | Только ЦП | ingest, job status, listener subscriptions |
| `/api/v1/*` | Внешние клиенты | Events с Bearer-ключом |
Internal защищён заголовком `X-Internal-Token` (`INTERNAL_TOKEN`).
### Профили, каналы и пары
1. **Канал** (`CRUD /admin/parse-channels`) — handle Telegram + метаданные.
2. **Профиль** (`CRUD /admin/parser-profiles`) — Generate (DeepSeek один раз по образцу) → Preview → Save `heuristic_profile`.
3. **Пара** (`POST /admin/jobs` с `channel_id` + `profile_id`) → `ParseJob`; при enqueue ЦА **разворачивает** пару в плоский `source_config`:
```json
{
"channel": "<ParseChannel.channel>",
"limit": 100,
"extract_mode": "profile",
"heuristic_profile": { "...из ParserProfile..." }
}
```
ЦП по-прежнему получает только Redis payload — без доступа к таблицам Profile/Channel.
Roadmap: кастомные пользовательские таблицы подменяют список `target-fields` при генерации; профиль будет ссылаться на `schema_id`.
### Jobs и scheduler
1. `POST /admin/jobs` / retry → запись `ParseJob` + `enqueue_parse_job` (flatten + RPUSH).
2. Очередь: `cp:jobs:{family}` из `contracts/queues.py`.
3. `services/scheduler.py` — тик ~30 с, повторная постановка активных jobs по `interval_seconds`.
### Ingest
`POST /internal/ingest` → `services/ingest.py`:
- дедуп по `source_url`;
- создание `Event`;
- при координатах — sync `MapObject`;
- batch обновляет статус job; `listener: true` — нет.
### Frontend (маршруты)
| Path | View |
|------|------|
| `/` | `MapViewPage.vue` |
| `/channels` | `ChannelsView.vue` |
| `/parser-profiles` | `ParserProfilesView.vue` |
| `/parsers` | `ParsersView.vue` (связки канал+профиль) |
| `/events` | `EventsView.vue` |
| `/analytics` | `AnalyticsView.vue` |
| `/consumers` | `ConsumersView.vue` |
Карта: Leaflet, фильтры дат/региона/темы/источника, CRUD объектов (ПКМ), медиа, таймлайн появления (если включён в UI).
### Nginx
`ca-frontend` слушает `:80` (с хоста `:8080`), проксирует backend-пути на `http://ca-api:8000`. Лимит тела для медиа задаётся в `nginx.conf`.
### Env (ЦА)
| Переменная | Назначение |
|------------|------------|
| `DATABASE_URL` | PostgreSQL |
| `REDIS_URL` | Очереди |
| `INTERNAL_TOKEN` | Auth ЦП ↔ ЦА |
| `TEST_PI_API_KEY` | Seed consumer `test-pi` |
| `UPLOAD_DIR` | Медиа (по умолчанию `/data/uploads`) |
| `DEEPSEEK_API_KEY` | Generate профиля на ca-api; не нужен для preview/runtime |
| `DEEPSEEK_BASE_URL` / `DEEPSEEK_MODEL` | Опционально |
### Типовые точки входа в код
| Задача | Файл |
|--------|------|
| Новый admin endpoint | `routers/admin.py` |
| Логика ingest | `services/ingest.py` |
| Фильтры карты | `services/map_query.py` + `routers/map.py` |
| Новый экран UI | `frontend/src/views/` + `router/index.ts` |
| Поля парсера в форме | `ParsersView.vue` (+ contracts) |
+91
View File
@@ -0,0 +1,91 @@
# Distribution API (ПИ)
**ПИ** — не отдельный Docker-сервис, а внешний HTTP-срез поверх данных ЦА.
Общий контекст: [docs/architecture-overview.md](../../docs/architecture-overview.md).
---
## Общими чертами
1. В UI «ПИ» (`/consumers`) создаётся подписчик.
2. Ему выдаётся API-ключ (в БД хранится только SHA-256 hash).
3. Клиент читает события:
```http
GET /api/v1/events
Authorization: Bearer <api_key>
```
4. ЦА отдаёт события с учётом фильтров consumer (регионы, темы, дата).
```mermaid
flowchart LR
Client[External system]
V1[GET /api/v1/events]
Cons[Consumer + Filter]
Events[(events)]
Client -->|Bearer key| V1
V1 --> Cons
Cons --> Events
```
---
## Подробнее
### Код
| Часть | Путь |
|-------|------|
| Роут | `centers/analytics/api/app/routers/v1.py` |
| Модели | `Consumer`, `ConsumerFilter` в `models.py` |
| Фильтрация | `services/filtering.py` |
| Admin CRUD | `routers/admin.py` — `/admin/consumers`, `…/rotate-key` |
| UI | `frontend/src/views/ConsumersView.vue` |
| Seed | `seed.py` — consumer `test-pi` из `TEST_PI_API_KEY` |
### Аутентификация
- Заголовок: `Authorization: Bearer <plain_api_key>`.
- Сравнение с `consumers.api_key_hash` (SHA-256).
- Неактивный consumer → отказ.
### Фильтры подписчика
Типичные ограничения (через `ConsumerFilter`):
- список регионов;
- список тем;
- `date_from` (нижняя граница даты события).
Точный набор полей — в модели/схемах admin API; при изменении фильтров обновляйте UI consumers и `filtering.py`.
### Admin операции
| Метод | Путь | Действие |
|-------|------|----------|
| GET | `/admin/consumers` | Список |
| POST | `/admin/consumers` | Создать (+ ключ в ответе один раз) |
| PATCH | `/admin/consumers/{id}` | Обновить |
| POST | `/admin/consumers/{id}/rotate-key` | Новый ключ |
### Локальный тест
После `docker compose up` (если не меняли seed):
```bash
curl -s -H "Authorization: Bearer test-pi-api-key-change-me" \
"http://localhost:8080/api/v1/events" | head
```
В проде обязательно смените `TEST_PI_API_KEY` и ротируйте ключи.
### Что ПИ не делает
- не пишет в БД;
- не ставит parse jobs;
- не ходит в Redis / ЦП.
Только чтение уже ingest'нутых событий через контракт `/api/v1`.
+74
View File
@@ -0,0 +1,74 @@
"""Admin auth: single env-based user + JWT bearer tokens."""
from __future__ import annotations
import hmac
import os
import time
from typing import Any
import jwt
from fastapi import HTTPException
ALGORITHM = "HS256"
DEFAULT_TTL_SECONDS = 60 * 60 * 24 # 24h
def _admin_user() -> str:
return os.getenv("ADMIN_USER", "admin").strip() or "admin"
def _admin_password() -> str:
return os.getenv("ADMIN_PASSWORD", "").strip()
def _jwt_secret() -> str:
secret = os.getenv("ADMIN_JWT_SECRET", "").strip()
if not secret:
# Dev fallback: derive from password so local stacks boot without extra secret.
password = _admin_password()
if not password:
raise HTTPException(
status_code=503,
detail="ADMIN_PASSWORD is not configured",
)
return f"mapmil-dev:{password}"
return secret
def admin_credentials_configured() -> bool:
return bool(_admin_password())
def verify_credentials(username: str, password: str) -> bool:
expected_user = _admin_user()
expected_password = _admin_password()
if not expected_password:
return False
user_ok = hmac.compare_digest(username.strip(), expected_user)
pass_ok = hmac.compare_digest(password, expected_password)
return user_ok and pass_ok
def create_access_token(*, username: str, ttl_seconds: int = DEFAULT_TTL_SECONDS) -> str:
now = int(time.time())
payload: dict[str, Any] = {
"sub": username,
"role": "admin",
"iat": now,
"exp": now + ttl_seconds,
}
return jwt.encode(payload, _jwt_secret(), algorithm=ALGORITHM)
def decode_access_token(token: str) -> dict[str, Any]:
try:
payload = jwt.decode(token, _jwt_secret(), algorithms=[ALGORITHM])
except jwt.ExpiredSignatureError as exc:
raise HTTPException(status_code=401, detail="Token expired") from exc
except jwt.InvalidTokenError as exc:
raise HTTPException(status_code=401, detail="Invalid token") from exc
if payload.get("role") != "admin" or not payload.get("sub"):
raise HTTPException(status_code=401, detail="Invalid token")
return payload
+15
View File
@@ -2,6 +2,8 @@ import os
from fastapi import Header, HTTPException
from .auth import decode_access_token
def verify_internal_token(
x_internal_token: str | None = Header(default=None, alias="X-Internal-Token"),
@@ -9,3 +11,16 @@ def verify_internal_token(
expected = os.getenv("INTERNAL_TOKEN", "dev-internal-token")
if not x_internal_token or x_internal_token != expected:
raise HTTPException(status_code=401, detail="Invalid internal token")
def verify_admin(
authorization: str | None = Header(default=None),
) -> str:
"""Require Authorization: Bearer <admin JWT>. Returns username (sub)."""
if not authorization or not authorization.startswith("Bearer "):
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
token = authorization.removeprefix("Bearer ").strip()
if not token:
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
payload = decode_access_token(token)
return str(payload["sub"])
+7 -2
View File
@@ -13,8 +13,8 @@ for _candidate in (_HERE.parent, *_HERE.parents):
break
from .database import Base, engine, get_db
from .routers import admin, internal, map, objects, v1
from .seed import seed_objects, seed_test_consumer
from .routers import admin, auth, internal, map, objects, parse_channels, parser_profiles, v1, vpn
from .seed import seed_objects, seed_test_consumer, seed_vpn_settings
from .services.migrations import migrate_schema
from .services.scheduler import start_scheduler
from .storage import ensure_upload_dir
@@ -30,6 +30,7 @@ async def lifespan(_: FastAPI):
try:
seed_objects(db)
seed_test_consumer(db)
seed_vpn_settings(db)
finally:
db.close()
yield
@@ -49,5 +50,9 @@ app.add_middleware(
app.include_router(objects.router)
app.include_router(map.router)
app.include_router(internal.router)
app.include_router(auth.router)
app.include_router(admin.router)
app.include_router(parser_profiles.router)
app.include_router(parse_channels.router)
app.include_router(vpn.router)
app.include_router(v1.router)
+67
View File
@@ -93,12 +93,57 @@ class Event(Base):
)
class ParserProfile(Base):
__tablename__ = "parser_profiles"
id: Mapped[int] = mapped_column(Integer, primary_key=True, index=True)
name: Mapped[str] = mapped_column(String(255), nullable=False)
# heuristic = static rules; llm = instruction + extract_schema at CP runtime
kind: Mapped[str] = mapped_column(String(50), default="heuristic", nullable=False, index=True)
sample_post: Mapped[str] = mapped_column(Text, default="")
heuristic_profile: Mapped[dict | None] = mapped_column(JSON, nullable=True)
llm_profile: Mapped[dict | None] = mapped_column(JSON, nullable=True)
status: Mapped[str] = mapped_column(String(50), default="draft", index=True)
created_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True),
default=lambda: datetime.now(timezone.utc),
)
jobs: Mapped[list["ParseJob"]] = relationship(back_populates="profile")
class ParseChannel(Base):
__tablename__ = "parse_channels"
id: Mapped[int] = mapped_column(Integer, primary_key=True, index=True)
name: Mapped[str] = mapped_column(String(255), nullable=False)
source_type: Mapped[str] = mapped_column(String(50), nullable=False, default="telegram")
channel: Mapped[str] = mapped_column(String(255), nullable=False)
is_active: Mapped[bool] = mapped_column(Boolean, default=True, nullable=False)
created_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True),
default=lambda: datetime.now(timezone.utc),
)
jobs: Mapped[list["ParseJob"]] = relationship(back_populates="channel")
class ParseJob(Base):
__tablename__ = "parse_jobs"
id: Mapped[int] = mapped_column(Integer, primary_key=True, index=True)
source_type: Mapped[str] = mapped_column(String(50), nullable=False)
source_config: Mapped[dict] = mapped_column(JSON, nullable=False, default=dict)
profile_id: Mapped[int | None] = mapped_column(
ForeignKey("parser_profiles.id", ondelete="SET NULL"),
nullable=True,
index=True,
)
channel_id: Mapped[int | None] = mapped_column(
ForeignKey("parse_channels.id", ondelete="SET NULL"),
nullable=True,
index=True,
)
schedule: Mapped[str | None] = mapped_column(String(100), nullable=True)
interval_seconds: Mapped[int] = mapped_column(Integer, default=3600, nullable=False)
is_active: Mapped[bool] = mapped_column(Boolean, default=True, nullable=False)
@@ -110,6 +155,9 @@ class ParseJob(Base):
default=lambda: datetime.now(timezone.utc),
)
profile: Mapped["ParserProfile | None"] = relationship(back_populates="jobs")
channel: Mapped["ParseChannel | None"] = relationship(back_populates="jobs")
class Consumer(Base):
__tablename__ = "consumers"
@@ -145,3 +193,22 @@ class ConsumerFilter(Base):
date_from: Mapped[datetime | None] = mapped_column(DateTime(timezone=True), nullable=True)
consumer: Mapped["Consumer"] = relationship(back_populates="filter")
class VpnSettingsRow(Base):
"""Singleton VPN / proxy settings (id=1)."""
__tablename__ = "vpn_settings"
id: Mapped[int] = mapped_column(Integer, primary_key=True)
enabled: Mapped[bool] = mapped_column(Boolean, default=False, nullable=False)
mode: Mapped[str] = mapped_column(String(32), default="subscription", nullable=False)
subscription_url: Mapped[str | None] = mapped_column(Text, nullable=True)
subscription_interval_seconds: Mapped[int] = mapped_column(
Integer, default=3600, nullable=False
)
host: Mapped[str | None] = mapped_column(String(255), nullable=True)
port: Mapped[int | None] = mapped_column(Integer, nullable=True)
username: Mapped[str | None] = mapped_column(String(255), nullable=True)
password: Mapped[str | None] = mapped_column(String(255), nullable=True)
proxied_source_types: Mapped[list | None] = mapped_column(JSON, nullable=True)
+235 -22
View File
@@ -4,14 +4,17 @@ from fastapi import APIRouter, Depends, HTTPException, Query
from sqlalchemy.orm import Session
from ..database import get_db
from ..models import Consumer, Event, ParseJob
from ..deps import verify_admin
from ..models import Consumer, Event, MapObject, ParseChannel, ParseJob, ParserProfile
from ..schemas import (
AnalyticsSummary,
ConsumerCreate,
ConsumerRead,
ConsumerUpdate,
EventCreate,
EventListResponse,
EventRead,
EventUpdate,
ParseJobCreate,
ParseJobRead,
ParseJobUpdate,
@@ -30,11 +33,12 @@ from ..services.filtering import (
create_consumer,
list_consumers,
rotate_consumer_key,
sync_event_to_map_object,
update_consumer,
)
from ..services.jobs import enqueue_job
from ..services.jobs import enqueue_parse_job, flatten_pair_config
router = APIRouter(prefix="/admin", tags=["admin"])
router = APIRouter(prefix="/admin", tags=["admin"], dependencies=[Depends(verify_admin)])
def _validated_source_config(source_type: str, source_config: dict) -> dict:
@@ -48,28 +52,89 @@ def _validated_source_config(source_type: str, source_config: dict) -> dict:
raise HTTPException(status_code=400, detail=str(exc)) from exc
def _job_to_read(job: ParseJob) -> ParseJobRead:
return ParseJobRead(
id=job.id,
source_type=job.source_type,
source_config=job.source_config or {},
profile_id=job.profile_id,
channel_id=job.channel_id,
schedule=job.schedule,
interval_seconds=job.interval_seconds,
is_active=job.is_active,
status=job.status,
last_run_at=job.last_run_at,
last_error=job.last_error,
created_at=job.created_at,
profile_name=job.profile.name if job.profile else None,
channel_name=job.channel.name if job.channel else None,
channel_handle=job.channel.channel if job.channel else None,
)
@router.post("/jobs", response_model=ParseJobRead, status_code=201)
def create_parse_job(payload: ParseJobCreate, db: Session = Depends(get_db)):
source_config = _validated_source_config(payload.source_type, payload.source_config)
job = ParseJob(
source_type=payload.source_type,
source_config=source_config,
schedule=payload.schedule,
interval_seconds=payload.interval_seconds,
is_active=payload.is_active,
status="queued",
)
pair_mode = payload.profile_id is not None or payload.channel_id is not None
if pair_mode:
if payload.profile_id is None or payload.channel_id is None:
raise HTTPException(
status_code=400,
detail="Both profile_id and channel_id are required to create a pair",
)
profile = db.query(ParserProfile).filter(ParserProfile.id == payload.profile_id).first()
channel = db.query(ParseChannel).filter(ParseChannel.id == payload.channel_id).first()
if not profile:
raise HTTPException(status_code=404, detail="Profile not found")
if not channel:
raise HTTPException(status_code=404, detail="Channel not found")
kind = (profile.kind or "heuristic").strip().lower()
if kind == "llm":
if not profile.llm_profile:
raise HTTPException(status_code=400, detail="LLM profile has no llm_profile")
elif not profile.heuristic_profile:
raise HTTPException(status_code=400, detail="Profile has no heuristic_profile")
if not channel.is_active:
raise HTTPException(status_code=400, detail="Channel is inactive")
try:
source_config = flatten_pair_config(channel, profile, limit=payload.limit)
except Exception as exc:
raise HTTPException(status_code=400, detail=str(exc)) from exc
job = ParseJob(
source_type=channel.source_type or "telegram",
source_config=source_config,
profile_id=profile.id,
channel_id=channel.id,
schedule=payload.schedule,
interval_seconds=payload.interval_seconds,
is_active=payload.is_active,
status="queued",
)
else:
source_config = _validated_source_config(payload.source_type, payload.source_config)
job = ParseJob(
source_type=payload.source_type,
source_config=source_config,
schedule=payload.schedule,
interval_seconds=payload.interval_seconds,
is_active=payload.is_active,
status="queued",
)
db.add(job)
db.commit()
db.refresh(job)
enqueue_job(job.id, job.source_type, job.source_config)
return job
enqueue_parse_job(db, job)
db.refresh(job)
return _job_to_read(job)
@router.get("/jobs", response_model=list[ParseJobRead])
def list_parse_jobs(db: Session = Depends(get_db)):
return db.query(ParseJob).order_by(ParseJob.id.desc()).all()
jobs = db.query(ParseJob).order_by(ParseJob.id.desc()).all()
return [_job_to_read(job) for job in jobs]
@router.get("/jobs/{job_id}", response_model=ParseJobRead)
@@ -77,7 +142,7 @@ def get_parse_job(job_id: int, db: Session = Depends(get_db)):
job = db.query(ParseJob).filter(ParseJob.id == job_id).first()
if not job:
raise HTTPException(status_code=404, detail="Job not found")
return job
return _job_to_read(job)
@router.post("/jobs/{job_id}/retry", response_model=ParseJobRead)
@@ -85,7 +150,10 @@ def retry_parse_job(job_id: int, db: Session = Depends(get_db)):
job = db.query(ParseJob).filter(ParseJob.id == job_id).first()
if not job:
raise HTTPException(status_code=404, detail="Job not found")
if job.status in ("queued", "running"):
from ..services.job_stale import is_stale_job
if job.status in ("queued", "running") and not is_stale_job(job):
raise HTTPException(status_code=409, detail="Job is already running or queued")
job.status = "queued"
@@ -93,8 +161,12 @@ def retry_parse_job(job_id: int, db: Session = Depends(get_db)):
db.commit()
db.refresh(job)
enqueue_job(job.id, job.source_type, job.source_config)
return job
try:
enqueue_parse_job(db, job)
except ValueError as exc:
raise HTTPException(status_code=400, detail=str(exc)) from exc
db.refresh(job)
return _job_to_read(job)
@router.patch("/jobs/{job_id}", response_model=ParseJobRead)
@@ -107,8 +179,36 @@ def update_parse_job(
if not job:
raise HTTPException(status_code=404, detail="Job not found")
if payload.source_config is not None:
if payload.profile_id is not None:
profile = db.query(ParserProfile).filter(ParserProfile.id == payload.profile_id).first()
if not profile:
raise HTTPException(status_code=404, detail="Profile not found")
job.profile_id = profile.id
if payload.channel_id is not None:
channel = db.query(ParseChannel).filter(ParseChannel.id == payload.channel_id).first()
if not channel:
raise HTTPException(status_code=404, detail="Channel not found")
job.channel_id = channel.id
if job.profile_id and job.channel_id:
profile = db.query(ParserProfile).filter(ParserProfile.id == job.profile_id).first()
channel = db.query(ParseChannel).filter(ParseChannel.id == job.channel_id).first()
if profile and channel:
existing_limit = (job.source_config or {}).get("limit", 100)
limit = payload.limit if payload.limit is not None else existing_limit
if not isinstance(limit, int):
limit = 100
try:
job.source_config = flatten_pair_config(channel, profile, limit=limit)
except Exception as exc:
raise HTTPException(status_code=400, detail=str(exc)) from exc
elif payload.source_config is not None:
job.source_config = _validated_source_config(job.source_type, payload.source_config)
elif payload.limit is not None and isinstance(job.source_config, dict):
cfg = dict(job.source_config)
cfg["limit"] = payload.limit
job.source_config = _validated_source_config(job.source_type, cfg)
if payload.interval_seconds is not None:
job.interval_seconds = payload.interval_seconds
if payload.is_active is not None:
@@ -116,7 +216,7 @@ def update_parse_job(
db.commit()
db.refresh(job)
return job
return _job_to_read(job)
@router.delete("/jobs/{job_id}", status_code=204)
@@ -125,7 +225,14 @@ def delete_parse_job(job_id: int, db: Session = Depends(get_db)):
if not job:
raise HTTPException(status_code=404, detail="Job not found")
if job.status == "running":
raise HTTPException(status_code=409, detail="Cannot delete a running job")
from ..services.job_stale import is_stale_job
if not is_stale_job(job):
raise HTTPException(status_code=409, detail="Cannot delete a running job")
# Stale running — allow delete after marking failed for audit trail
job.status = "failed"
job.last_error = "Deleted while stale running"
db.commit()
db.delete(job)
db.commit()
@@ -159,6 +266,112 @@ def list_events(
return EventListResponse(items=items, total=total)
def _ensure_unique_source_url(db: Session, source_url: str, *, exclude_id: int | None = None) -> None:
query = db.query(Event).filter(Event.source_url == source_url)
if exclude_id is not None:
query = query.filter(Event.id != exclude_id)
if query.first():
raise HTTPException(status_code=409, detail="Event with this source_url already exists")
def _sync_or_clear_map_object(db: Session, event: Event) -> None:
if event.latitude is not None and event.longitude is not None:
sync_event_to_map_object(db, event)
return
linked = db.query(MapObject).filter(MapObject.event_id == event.id).first()
if linked:
db.delete(linked)
@router.post("/events", response_model=EventRead, status_code=201)
def create_event(payload: EventCreate, db: Session = Depends(get_db)):
source_url = payload.source_url.strip()
if not source_url:
raise HTTPException(status_code=400, detail="source_url is required")
_ensure_unique_source_url(db, source_url)
if (payload.latitude is None) != (payload.longitude is None):
raise HTTPException(status_code=400, detail="latitude and longitude must be set together")
event = Event(
source_type=payload.source_type.strip() or "manual",
source_url=source_url,
raw_text=payload.raw_text or "",
title=payload.title or "",
description=payload.description or "",
locality=payload.locality or "",
latitude=payload.latitude,
longitude=payload.longitude,
event_date=payload.event_date,
region=payload.region,
topic=payload.topic,
tags=payload.tags,
metadata_=payload.metadata,
)
db.add(event)
db.flush()
_sync_or_clear_map_object(db, event)
db.commit()
db.refresh(event)
return event
@router.get("/events/{event_id}", response_model=EventRead)
def get_event(event_id: int, db: Session = Depends(get_db)):
event = db.query(Event).filter(Event.id == event_id).first()
if not event:
raise HTTPException(status_code=404, detail="Event not found")
return event
@router.patch("/events/{event_id}", response_model=EventRead)
def update_event(event_id: int, payload: EventUpdate, db: Session = Depends(get_db)):
event = db.query(Event).filter(Event.id == event_id).first()
if not event:
raise HTTPException(status_code=404, detail="Event not found")
data = payload.model_dump(exclude_unset=True)
if "source_url" in data and data["source_url"] is not None:
source_url = data["source_url"].strip()
if not source_url:
raise HTTPException(status_code=400, detail="source_url cannot be empty")
_ensure_unique_source_url(db, source_url, exclude_id=event.id)
event.source_url = source_url
data.pop("source_url")
if "metadata" in data:
event.metadata_ = data.pop("metadata")
if "source_type" in data and data["source_type"] is not None:
event.source_type = data.pop("source_type").strip() or event.source_type
for field, value in data.items():
setattr(event, field, value)
lat = event.latitude
lon = event.longitude
if (lat is None) != (lon is None):
raise HTTPException(status_code=400, detail="latitude and longitude must be set together")
_sync_or_clear_map_object(db, event)
db.commit()
db.refresh(event)
return event
@router.delete("/events/{event_id}", status_code=204)
def delete_event(event_id: int, db: Session = Depends(get_db)):
event = db.query(Event).filter(Event.id == event_id).first()
if not event:
raise HTTPException(status_code=404, detail="Event not found")
linked = db.query(MapObject).filter(MapObject.event_id == event.id).all()
for obj in linked:
db.delete(obj)
db.delete(event)
db.commit()
@router.get("/analytics/summary", response_model=AnalyticsSummary)
def analytics_summary(db: Session = Depends(get_db)):
return get_analytics_summary(db)
+41
View File
@@ -0,0 +1,41 @@
from fastapi import APIRouter, Depends, HTTPException
from pydantic import BaseModel, Field
from ..auth import (
admin_credentials_configured,
create_access_token,
verify_credentials,
)
from ..deps import verify_admin
router = APIRouter(prefix="/admin/auth", tags=["auth"])
class LoginRequest(BaseModel):
username: str = Field(min_length=1)
password: str = Field(min_length=1)
class TokenResponse(BaseModel):
access_token: str
token_type: str = "bearer"
class MeResponse(BaseModel):
username: str
role: str = "admin"
@router.post("/login", response_model=TokenResponse)
def login(payload: LoginRequest):
if not admin_credentials_configured():
raise HTTPException(status_code=503, detail="Admin auth is not configured")
if not verify_credentials(payload.username, payload.password):
raise HTTPException(status_code=401, detail="Invalid username or password")
token = create_access_token(username=payload.username.strip())
return TokenResponse(access_token=token)
@router.get("/me", response_model=MeResponse)
def me(username: str = Depends(verify_admin)):
return MeResponse(username=username)
+32 -4
View File
@@ -6,8 +6,15 @@ from sqlalchemy.orm import Session
from ..database import get_db
from ..deps import verify_internal_token
from ..models import ParseJob
from ..schemas import IngestRequest, IngestResponse, ListenerSubscription
from ..schemas import (
IngestRequest,
IngestResponse,
ListenerSubscription,
VpnSettingsInternalRead,
)
from ..services.ingest import ingest_events
from ..services.jobs import resolve_job_source_config
from ..services.vpn import ensure_vpn_settings, to_internal_read
router = APIRouter(prefix="/internal", tags=["internal"])
@@ -45,7 +52,8 @@ def update_job_status(
raise HTTPException(status_code=404, detail="Job not found")
job.status = status
if status in ("completed", "failed"):
# Anchor staleness detection: running/queued start, and terminal finish
if status in ("running", "queued", "completed", "failed"):
job.last_run_at = datetime.now(timezone.utc)
job.last_error = error
db.commit()
@@ -68,8 +76,28 @@ def listener_subscriptions(
)
result: list[ListenerSubscription] = []
for job in jobs:
channel = job.source_config.get("channel") if job.source_config else None
try:
source_config = resolve_job_source_config(db, job)
except ValueError:
continue
channel = source_config.get("channel")
if not channel:
continue
result.append(ListenerSubscription(job_id=job.id, channel=str(channel)))
result.append(
ListenerSubscription(
job_id=job.id,
channel=str(channel),
source_config=dict(source_config),
)
)
db.commit()
return result
@router.get("/vpn", response_model=VpnSettingsInternalRead)
def internal_vpn_settings(
_: None = Depends(verify_internal_token),
db: Session = Depends(get_db),
):
row = ensure_vpn_settings(db)
return to_internal_read(row)
+18 -3
View File
@@ -3,6 +3,7 @@ from fastapi.responses import FileResponse
from sqlalchemy.orm import Session
from ..database import get_db
from ..deps import verify_admin
from ..models import MapObject, ObjectMedia
from ..schemas import MapObjectCreate, MapObjectRead, MapObjectUpdate, ObjectMediaRead
from ..storage import (
@@ -48,7 +49,11 @@ def get_object(object_id: int, db: Session = Depends(get_db)):
@router.post("/api/objects", response_model=MapObjectRead, status_code=201)
def create_object(payload: MapObjectCreate, db: Session = Depends(get_db)):
def create_object(
payload: MapObjectCreate,
db: Session = Depends(get_db),
_: str = Depends(verify_admin),
):
data = payload.model_dump()
created_at = data.pop("created_at", None)
obj = MapObject(**data)
@@ -65,6 +70,7 @@ def update_object(
object_id: int,
payload: MapObjectUpdate,
db: Session = Depends(get_db),
_: str = Depends(verify_admin),
):
obj = db.query(MapObject).filter(MapObject.id == object_id).first()
if not obj:
@@ -79,7 +85,11 @@ def update_object(
@router.delete("/api/objects/{object_id}", status_code=204)
def delete_object(object_id: int, db: Session = Depends(get_db)):
def delete_object(
object_id: int,
db: Session = Depends(get_db),
_: str = Depends(verify_admin),
):
obj = db.query(MapObject).filter(MapObject.id == object_id).first()
if not obj:
raise HTTPException(status_code=404, detail="Объект не найден")
@@ -111,6 +121,7 @@ async def upload_object_media(
object_id: int,
file: UploadFile = File(...),
db: Session = Depends(get_db),
_: str = Depends(verify_admin),
):
obj = db.query(MapObject).filter(MapObject.id == object_id).first()
if not obj:
@@ -159,7 +170,11 @@ def get_media_file(media_id: int, db: Session = Depends(get_db)):
@router.delete("/api/media/{media_id}", status_code=204)
def delete_media(media_id: int, db: Session = Depends(get_db)):
def delete_media(
media_id: int,
db: Session = Depends(get_db),
_: str = Depends(verify_admin),
):
media = db.query(ObjectMedia).filter(ObjectMedia.id == media_id).first()
if not media:
raise HTTPException(status_code=404, detail="Медиафайл не найден")
@@ -0,0 +1,103 @@
"""Admin CRUD for parse channels (Telegram handles in MVP)."""
from __future__ import annotations
from fastapi import APIRouter, Depends, HTTPException
from sqlalchemy.orm import Session
from ..database import get_db
from ..deps import verify_admin
from ..models import ParseChannel, ParseJob
from ..schemas import ParseChannelCreate, ParseChannelRead, ParseChannelUpdate
router = APIRouter(
prefix="/admin/parse-channels",
tags=["parse-channels"],
dependencies=[Depends(verify_admin)],
)
ALLOWED_SOURCE_TYPES = {"telegram"}
@router.get("", response_model=list[ParseChannelRead])
def list_channels(db: Session = Depends(get_db)):
return db.query(ParseChannel).order_by(ParseChannel.id.desc()).all()
@router.post("", response_model=ParseChannelRead, status_code=201)
def create_channel(payload: ParseChannelCreate, db: Session = Depends(get_db)):
source_type = (payload.source_type or "telegram").strip()
if source_type not in ALLOWED_SOURCE_TYPES:
raise HTTPException(
status_code=400,
detail=f"source_type must be one of: {', '.join(sorted(ALLOWED_SOURCE_TYPES))}",
)
channel = ParseChannel(
name=payload.name.strip(),
source_type=source_type,
channel=payload.channel.strip().lstrip("@"),
is_active=payload.is_active,
)
db.add(channel)
db.commit()
db.refresh(channel)
return channel
@router.get("/{channel_id}", response_model=ParseChannelRead)
def get_channel(channel_id: int, db: Session = Depends(get_db)):
channel = db.query(ParseChannel).filter(ParseChannel.id == channel_id).first()
if not channel:
raise HTTPException(status_code=404, detail="Channel not found")
return channel
@router.patch("/{channel_id}", response_model=ParseChannelRead)
def update_channel(
channel_id: int,
payload: ParseChannelUpdate,
db: Session = Depends(get_db),
):
channel = db.query(ParseChannel).filter(ParseChannel.id == channel_id).first()
if not channel:
raise HTTPException(status_code=404, detail="Channel not found")
if payload.name is not None:
channel.name = payload.name.strip()
if payload.source_type is not None:
source_type = payload.source_type.strip()
if source_type not in ALLOWED_SOURCE_TYPES:
raise HTTPException(
status_code=400,
detail=f"source_type must be one of: {', '.join(sorted(ALLOWED_SOURCE_TYPES))}",
)
channel.source_type = source_type
if payload.channel is not None:
channel.channel = payload.channel.strip().lstrip("@")
if payload.is_active is not None:
channel.is_active = payload.is_active
db.commit()
db.refresh(channel)
return channel
@router.delete("/{channel_id}", status_code=204)
def delete_channel(channel_id: int, db: Session = Depends(get_db)):
channel = db.query(ParseChannel).filter(ParseChannel.id == channel_id).first()
if not channel:
raise HTTPException(status_code=404, detail="Channel not found")
linked = (
db.query(ParseJob)
.filter(ParseJob.channel_id == channel_id, ParseJob.status.in_(("queued", "running")))
.count()
)
if linked:
raise HTTPException(
status_code=409,
detail="Cannot delete channel used by queued/running jobs",
)
db.delete(channel)
db.commit()
@@ -0,0 +1,304 @@
"""Admin CRUD for reusable parser profiles + generate/preview."""
from __future__ import annotations
from typing import Any
from fastapi import APIRouter, Depends, HTTPException
from pydantic import BaseModel, Field
from sqlalchemy.orm import Session
from contracts.heuristic_profile import HeuristicProfile
from contracts.llm_profile import LlmProfile
from ..database import get_db
from ..deps import verify_admin
from ..models import ParseJob, ParserProfile
from ..schemas import ParserProfileCreate, ParserProfileRead, ParserProfileUpdate
from ..services import parser_builder as builder
router = APIRouter(
prefix="/admin/parser-profiles",
tags=["parser-profiles"],
dependencies=[Depends(verify_admin)],
)
class GenerateRequest(BaseModel):
sample_post: str = Field(min_length=1)
hint: str | None = None
current_profile: dict[str, Any] | None = None
class GenerateResponse(BaseModel):
profile: dict[str, Any]
preview: dict[str, str]
empty_fields: list[str] = Field(default_factory=list)
class PreviewRequest(BaseModel):
sample_post: str = Field(min_length=1)
heuristic_profile: dict[str, Any] | None = None
# Accept alias used by older builder UI
profile: dict[str, Any] | None = None
class PreviewResponse(BaseModel):
fields: dict[str, str]
matched: bool = True
missing_required: list[str] = Field(default_factory=list)
class PreviewLlmRequest(BaseModel):
sample_post: str = Field(min_length=1)
llm_profile: dict[str, Any]
class PreviewLlmResponse(BaseModel):
fields: dict[str, str]
events: list[dict[str, str]] = Field(default_factory=list)
matched: bool = True
missing_required: list[str] = Field(default_factory=list)
is_event: bool = True
matched_count: int = 0
def _validate_heuristic_profile(raw: dict[str, Any] | None) -> dict[str, Any] | None:
if raw is None:
return None
try:
return HeuristicProfile.model_validate(raw).model_dump()
except Exception as exc:
raise HTTPException(status_code=400, detail=f"Invalid heuristic_profile: {exc}") from exc
def _validate_llm_profile(raw: dict[str, Any] | None) -> dict[str, Any] | None:
if raw is None:
return None
try:
return LlmProfile.model_validate(raw).model_dump()
except Exception as exc:
raise HTTPException(status_code=400, detail=f"Invalid llm_profile: {exc}") from exc
def _normalize_kind(kind: str | None) -> str:
value = (kind or "heuristic").strip().lower()
if value not in ("heuristic", "llm"):
raise HTTPException(status_code=400, detail="kind must be heuristic or llm")
return value
def _profile_status(
*,
kind: str,
heuristic_profile: dict | None,
llm_profile: dict | None,
explicit: str | None = None,
) -> str:
if explicit:
return explicit
if kind == "llm":
return "ready" if llm_profile else "draft"
return "ready" if heuristic_profile else "draft"
@router.get("/target-fields")
def target_fields():
return {"fields": builder.get_target_fields()}
@router.get("/llm-defaults")
def llm_defaults():
from contracts.llm_profile import DEFAULT_EXTRACT_SCHEMA, DEFAULT_INSTRUCTION
return {
"instruction": DEFAULT_INSTRUCTION,
"extract_schema": DEFAULT_EXTRACT_SCHEMA,
"required_fields": [],
"multi_event": False,
}
@router.post("/generate", response_model=GenerateResponse)
async def generate_profile(payload: GenerateRequest):
if not builder.deepseek_enabled():
raise HTTPException(
status_code=503,
detail="DEEPSEEK_API_KEY is not set on ca-api. Add it to .env for parser generation.",
)
try:
profile = await builder.generate_profile(
payload.sample_post,
hint=payload.hint,
current_profile=payload.current_profile,
)
except ValueError as exc:
raise HTTPException(status_code=400, detail=str(exc)) from exc
except Exception as exc:
raise HTTPException(
status_code=502,
detail=f"DeepSeek generate failed: {exc}",
) from exc
dumped = profile.model_dump()
if payload.current_profile and isinstance(payload.current_profile.get("required_fields"), list):
dumped["required_fields"] = payload.current_profile["required_fields"]
dumped = HeuristicProfile.model_validate(dumped).model_dump()
preview = builder.preview_with_profile(payload.sample_post, dumped)
return GenerateResponse(
profile=dumped,
preview=preview,
empty_fields=builder.empty_preview_fields(preview),
)
@router.post("/preview", response_model=PreviewResponse)
def preview_profile(payload: PreviewRequest):
raw = payload.heuristic_profile if payload.heuristic_profile is not None else payload.profile
if raw is None:
raise HTTPException(status_code=400, detail="heuristic_profile is required")
try:
HeuristicProfile.model_validate(raw)
fields, matched, missing = builder.match_preview(payload.sample_post, raw)
except Exception as exc:
raise HTTPException(status_code=400, detail=str(exc)) from exc
return PreviewResponse(fields=fields, matched=matched, missing_required=missing)
@router.post("/preview-llm", response_model=PreviewLlmResponse)
async def preview_llm_profile(payload: PreviewLlmRequest):
if not builder.deepseek_enabled():
raise HTTPException(
status_code=503,
detail="DEEPSEEK_API_KEY is not set on ca-api. Add it to .env for LLM preview.",
)
try:
LlmProfile.model_validate(payload.llm_profile)
events, fields, matched, missing, is_event, matched_count = await builder.preview_llm_extract(
payload.sample_post,
payload.llm_profile,
)
except ValueError as exc:
raise HTTPException(status_code=400, detail=str(exc)) from exc
except Exception as exc:
raise HTTPException(
status_code=502,
detail=f"DeepSeek LLM preview failed: {exc}",
) from exc
return PreviewLlmResponse(
fields=fields,
events=events,
matched=matched,
missing_required=missing,
is_event=is_event,
matched_count=matched_count,
)
@router.get("", response_model=list[ParserProfileRead])
def list_profiles(db: Session = Depends(get_db)):
return db.query(ParserProfile).order_by(ParserProfile.id.desc()).all()
@router.post("", response_model=ParserProfileRead, status_code=201)
def create_profile(payload: ParserProfileCreate, db: Session = Depends(get_db)):
kind = _normalize_kind(payload.kind)
heuristic = _validate_heuristic_profile(payload.heuristic_profile)
llm = _validate_llm_profile(payload.llm_profile)
if kind == "heuristic" and llm and not heuristic:
# ignore stray llm blob when creating heuristic
llm = None
if kind == "llm" and heuristic and not llm:
heuristic = None
profile = ParserProfile(
name=payload.name.strip(),
kind=kind,
sample_post=payload.sample_post or "",
heuristic_profile=heuristic if kind == "heuristic" else None,
llm_profile=llm if kind == "llm" else None,
status=_profile_status(
kind=kind,
heuristic_profile=heuristic if kind == "heuristic" else None,
llm_profile=llm if kind == "llm" else None,
explicit=payload.status,
),
)
db.add(profile)
db.commit()
db.refresh(profile)
return profile
@router.get("/{profile_id}", response_model=ParserProfileRead)
def get_profile(profile_id: int, db: Session = Depends(get_db)):
profile = db.query(ParserProfile).filter(ParserProfile.id == profile_id).first()
if not profile:
raise HTTPException(status_code=404, detail="Profile not found")
return profile
@router.patch("/{profile_id}", response_model=ParserProfileRead)
def update_profile(
profile_id: int,
payload: ParserProfileUpdate,
db: Session = Depends(get_db),
):
profile = db.query(ParserProfile).filter(ParserProfile.id == profile_id).first()
if not profile:
raise HTTPException(status_code=404, detail="Profile not found")
if payload.name is not None:
profile.name = payload.name.strip()
if payload.sample_post is not None:
profile.sample_post = payload.sample_post
if payload.kind is not None:
profile.kind = _normalize_kind(payload.kind)
kind = _normalize_kind(profile.kind)
if "heuristic_profile" in payload.model_fields_set:
profile.heuristic_profile = _validate_heuristic_profile(payload.heuristic_profile)
if "llm_profile" in payload.model_fields_set:
profile.llm_profile = _validate_llm_profile(payload.llm_profile)
# Keep only the blob matching kind
if kind == "heuristic":
profile.llm_profile = None
else:
profile.heuristic_profile = None
if payload.status is not None:
profile.status = payload.status
elif (
"heuristic_profile" in payload.model_fields_set
or "llm_profile" in payload.model_fields_set
or payload.kind is not None
):
profile.status = _profile_status(
kind=kind,
heuristic_profile=profile.heuristic_profile,
llm_profile=profile.llm_profile,
)
db.commit()
db.refresh(profile)
return profile
@router.delete("/{profile_id}", status_code=204)
def delete_profile(profile_id: int, db: Session = Depends(get_db)):
profile = db.query(ParserProfile).filter(ParserProfile.id == profile_id).first()
if not profile:
raise HTTPException(status_code=404, detail="Profile not found")
linked = (
db.query(ParseJob)
.filter(ParseJob.profile_id == profile_id, ParseJob.status.in_(("queued", "running")))
.count()
)
if linked:
raise HTTPException(
status_code=409,
detail="Cannot delete profile used by queued/running jobs",
)
db.delete(profile)
db.commit()
+32
View File
@@ -0,0 +1,32 @@
"""Admin VPN / proxy settings."""
from __future__ import annotations
from fastapi import APIRouter, Depends, HTTPException
from sqlalchemy.orm import Session
from ..database import get_db
from ..deps import verify_admin
from ..schemas import VpnSettingsAdminRead, VpnSettingsUpdate
from ..services.vpn import apply_vpn_update, ensure_vpn_settings, to_admin_read
router = APIRouter(
prefix="/admin/vpn",
tags=["vpn"],
dependencies=[Depends(verify_admin)],
)
@router.get("", response_model=VpnSettingsAdminRead)
def get_vpn_settings(db: Session = Depends(get_db)):
row = ensure_vpn_settings(db)
return to_admin_read(row)
@router.put("", response_model=VpnSettingsAdminRead)
def put_vpn_settings(payload: VpnSettingsUpdate, db: Session = Depends(get_db)):
try:
row = apply_vpn_update(db, payload)
except ValueError as exc:
raise HTTPException(status_code=400, detail=str(exc)) from exc
return to_admin_read(row)
+140
View File
@@ -92,6 +92,38 @@ class EventRead(BaseModel):
metadata: dict[str, Any] | None = Field(validation_alias="metadata_")
class EventCreate(BaseModel):
source_type: str = Field(default="manual", min_length=1, max_length=50)
source_url: str = Field(min_length=1, max_length=512)
raw_text: str = ""
title: str = ""
description: str = ""
locality: str = ""
latitude: float | None = Field(default=None, ge=-90, le=90)
longitude: float | None = Field(default=None, ge=-180, le=180)
event_date: datetime | None = None
region: str | None = None
topic: str | None = None
tags: list[str] | None = None
metadata: dict[str, Any] | None = None
class EventUpdate(BaseModel):
source_type: str | None = Field(default=None, min_length=1, max_length=50)
source_url: str | None = Field(default=None, min_length=1, max_length=512)
raw_text: str | None = None
title: str | None = None
description: str | None = None
locality: str | None = None
latitude: float | None = Field(default=None, ge=-90, le=90)
longitude: float | None = Field(default=None, ge=-180, le=180)
event_date: datetime | None = None
region: str | None = None
topic: str | None = None
tags: list[str] | None = None
metadata: dict[str, Any] | None = None
class IngestEventItem(BaseModel):
source_type: str = "telegram"
source_url: str
@@ -117,6 +149,7 @@ class IngestRequest(BaseModel):
class ListenerSubscription(BaseModel):
job_id: int
channel: str
source_config: dict[str, Any] = Field(default_factory=dict)
class IngestResponse(BaseModel):
@@ -126,9 +159,68 @@ class IngestResponse(BaseModel):
map_objects_synced: int
class ParserProfileCreate(BaseModel):
name: str = Field(min_length=1, max_length=255)
kind: str = Field(default="heuristic", pattern="^(heuristic|llm)$")
sample_post: str = ""
heuristic_profile: dict[str, Any] | None = None
llm_profile: dict[str, Any] | None = None
status: str | None = None
class ParserProfileUpdate(BaseModel):
name: str | None = Field(default=None, min_length=1, max_length=255)
kind: str | None = Field(default=None, pattern="^(heuristic|llm)$")
sample_post: str | None = None
heuristic_profile: dict[str, Any] | None = None
llm_profile: dict[str, Any] | None = None
status: str | None = None
class ParserProfileRead(BaseModel):
model_config = ConfigDict(from_attributes=True)
id: int
name: str
kind: str = "heuristic"
sample_post: str
heuristic_profile: dict[str, Any] | None
llm_profile: dict[str, Any] | None = None
status: str
created_at: datetime
class ParseChannelCreate(BaseModel):
name: str = Field(min_length=1, max_length=255)
source_type: str = "telegram"
channel: str = Field(min_length=1, max_length=255)
is_active: bool = True
class ParseChannelUpdate(BaseModel):
name: str | None = Field(default=None, min_length=1, max_length=255)
source_type: str | None = None
channel: str | None = Field(default=None, min_length=1, max_length=255)
is_active: bool | None = None
class ParseChannelRead(BaseModel):
model_config = ConfigDict(from_attributes=True)
id: int
name: str
source_type: str
channel: str
is_active: bool
created_at: datetime
class ParseJobCreate(BaseModel):
source_type: str = "telegram"
source_config: dict[str, Any] = Field(default_factory=dict)
profile_id: int | None = None
channel_id: int | None = None
limit: int = Field(default=100, ge=1, le=1000)
schedule: str | None = None
interval_seconds: int = Field(default=3600, ge=60, le=604800)
is_active: bool = True
@@ -136,6 +228,9 @@ class ParseJobCreate(BaseModel):
class ParseJobUpdate(BaseModel):
source_config: dict[str, Any] | None = None
profile_id: int | None = None
channel_id: int | None = None
limit: int | None = Field(default=None, ge=1, le=1000)
interval_seconds: int | None = Field(default=None, ge=60, le=604800)
is_active: bool | None = None
@@ -146,6 +241,8 @@ class ParseJobRead(BaseModel):
id: int
source_type: str
source_config: dict[str, Any]
profile_id: int | None = None
channel_id: int | None = None
schedule: str | None
interval_seconds: int
is_active: bool
@@ -153,6 +250,9 @@ class ParseJobRead(BaseModel):
last_run_at: datetime | None
last_error: str | None
created_at: datetime
profile_name: str | None = None
channel_name: str | None = None
channel_handle: str | None = None
class ConsumerCreate(BaseModel):
@@ -206,3 +306,43 @@ class TimelinePoint(BaseModel):
class TopItem(BaseModel):
name: str
count: int
class VpnSettingsUpdate(BaseModel):
enabled: bool | None = None
mode: Literal["subscription", "socks5", "http"] | None = None
subscription_url: str | None = None
clear_subscription_url: bool | None = None
subscription_interval_seconds: int | None = Field(default=None, ge=60, le=86400)
host: str | None = None
port: int | None = Field(default=None, ge=1, le=65535)
username: str | None = None
password: str | None = None
clear_password: bool | None = None
proxied_source_types: list[str] | None = None
class VpnSettingsAdminRead(BaseModel):
enabled: bool
mode: Literal["subscription", "socks5", "http"]
subscription_url_set: bool
subscription_url_hint: str | None = None
subscription_interval_seconds: int
host: str | None = None
port: int | None = None
username: str | None = None
password_set: bool
proxied_source_types: list[str]
available_source_types: list[str]
class VpnSettingsInternalRead(BaseModel):
enabled: bool
mode: Literal["subscription", "socks5", "http"]
subscription_url: str | None = None
subscription_interval_seconds: int
host: str | None = None
port: int | None = None
username: str | None = None
password: str | None = None
proxied_source_types: list[str]
+5
View File
@@ -4,6 +4,7 @@ from sqlalchemy.orm import Session
from .models import Consumer, ConsumerFilter
from .services.filtering import hash_api_key
from .services.vpn import ensure_vpn_settings
def seed_objects(db: Session) -> None:
@@ -24,3 +25,7 @@ def seed_test_consumer(db: Session) -> None:
db.flush()
db.add(ConsumerFilter(consumer_id=consumer.id, regions=None, topics=None))
db.commit()
def seed_vpn_settings(db: Session) -> None:
ensure_vpn_settings(db)
@@ -0,0 +1,36 @@
"""Stale parse-job recovery helpers (running/queued left behind after worker crash)."""
from __future__ import annotations
import os
from datetime import datetime, timezone
from ..models import ParseJob
# Jobs stuck in running/queued longer than this are considered abandoned.
STALE_JOB_SECONDS = int(os.getenv("STALE_JOB_SECONDS", "900"))
def _aware(dt: datetime | None) -> datetime | None:
if dt is None:
return None
if dt.tzinfo is None:
return dt.replace(tzinfo=timezone.utc)
return dt
def job_anchor_time(job: ParseJob) -> datetime | None:
"""Best available timestamp for staleness (prefer last_run_at)."""
return _aware(job.last_run_at) or _aware(getattr(job, "created_at", None))
def is_stale_job(job: ParseJob, now: datetime | None = None, *, ttl: int | None = None) -> bool:
if job.status not in ("running", "queued"):
return False
now = now or datetime.now(timezone.utc)
anchor = job_anchor_time(job)
if anchor is None:
# No timestamp — treat long-lived running as stale immediately for recovery
return job.status == "running"
limit = ttl if ttl is not None else STALE_JOB_SECONDS
return (now - anchor).total_seconds() >= limit
+71 -1
View File
@@ -4,6 +4,7 @@ import sys
from pathlib import Path
import redis
from sqlalchemy.orm import Session
# Allow importing shared contracts from monorepo root in local runs
_HERE = Path(__file__).resolve()
@@ -14,7 +15,9 @@ for _candidate in (_HERE.parent, *_HERE.parents):
break
from contracts.queues import queue_key_for_source # noqa: E402
from contracts.sources import parse_source_config # noqa: E402
from contracts.sources import TelegramSourceConfig, parse_source_config # noqa: E402
from ..models import ParseChannel, ParseJob, ParserProfile # noqa: E402
REDIS_URL = os.getenv("REDIS_URL", "redis://redis:6379/0")
@@ -23,6 +26,66 @@ def get_redis() -> redis.Redis:
return redis.from_url(REDIS_URL, decode_responses=True)
def flatten_pair_config(
channel: ParseChannel,
profile: ParserProfile,
*,
limit: int = 100,
) -> dict:
"""Expand Profile + Channel into Redis/CP source_config."""
kind = (profile.kind or "heuristic").strip().lower()
if kind == "llm":
from contracts.llm_profile import LlmProfile
if not profile.llm_profile:
raise ValueError("ParserProfile.llm_profile is empty")
llm = LlmProfile.model_validate(profile.llm_profile)
cfg = TelegramSourceConfig(
channel=channel.channel.strip(),
limit=limit,
extract_mode="llm",
extract_schema=llm.extract_schema,
instruction=llm.instruction,
required_fields=list(llm.required_fields),
multi_event=bool(llm.multi_event),
sample_post=profile.sample_post or None,
)
return cfg.model_dump()
if not profile.heuristic_profile:
raise ValueError("ParserProfile.heuristic_profile is empty")
cfg = TelegramSourceConfig(
channel=channel.channel.strip(),
limit=limit,
extract_mode="profile",
heuristic_profile=profile.heuristic_profile,
sample_post=profile.sample_post or None,
)
return cfg.model_dump()
def resolve_job_source_config(db: Session, job: ParseJob) -> dict:
"""
For pair jobs, rebuild flat source_config from current Profile/Channel.
Legacy jobs keep stored source_config.
"""
if job.profile_id and job.channel_id:
profile = db.query(ParserProfile).filter(ParserProfile.id == job.profile_id).first()
channel = db.query(ParseChannel).filter(ParseChannel.id == job.channel_id).first()
if not profile:
raise ValueError(f"ParserProfile {job.profile_id} not found")
if not channel:
raise ValueError(f"ParseChannel {job.channel_id} not found")
existing = job.source_config or {}
limit = existing.get("limit", 100)
if not isinstance(limit, int) or limit < 1:
limit = 100
flattened = flatten_pair_config(channel, profile, limit=limit)
job.source_config = flattened
return flattened
return dict(job.source_config or {})
def enqueue_job(job_id: int, source_type: str, source_config: dict) -> None:
# Validate known configs early; unknown types still raise from queue_key_for_source
try:
@@ -38,3 +101,10 @@ def enqueue_job(job_id: int, source_type: str, source_config: dict) -> None:
}
key = queue_key_for_source(source_type)
get_redis().rpush(key, json.dumps(payload))
def enqueue_parse_job(db: Session, job: ParseJob) -> None:
"""Resolve (flatten if pair) then push to Redis."""
source_config = resolve_job_source_config(db, job)
db.commit()
enqueue_job(job.id, job.source_type, source_config)
@@ -5,20 +5,32 @@ from sqlalchemy.engine import Engine
def migrate_schema(engine: Engine) -> None:
"""Apply lightweight schema updates for existing deployments."""
inspector = inspect(engine)
if "parse_jobs" not in inspector.get_table_names():
return
columns = {col["name"] for col in inspector.get_columns("parse_jobs")}
tables = set(inspector.get_table_names())
statements: list[str] = []
if "interval_seconds" not in columns:
statements.append(
"ALTER TABLE parse_jobs ADD COLUMN interval_seconds INTEGER NOT NULL DEFAULT 3600"
)
if "is_active" not in columns:
statements.append(
"ALTER TABLE parse_jobs ADD COLUMN is_active BOOLEAN NOT NULL DEFAULT TRUE"
)
if "parse_jobs" in tables:
columns = {col["name"] for col in inspector.get_columns("parse_jobs")}
if "interval_seconds" not in columns:
statements.append(
"ALTER TABLE parse_jobs ADD COLUMN interval_seconds INTEGER NOT NULL DEFAULT 3600"
)
if "is_active" not in columns:
statements.append(
"ALTER TABLE parse_jobs ADD COLUMN is_active BOOLEAN NOT NULL DEFAULT TRUE"
)
if "profile_id" not in columns:
statements.append("ALTER TABLE parse_jobs ADD COLUMN profile_id INTEGER")
if "channel_id" not in columns:
statements.append("ALTER TABLE parse_jobs ADD COLUMN channel_id INTEGER")
if "parser_profiles" in tables:
columns = {col["name"] for col in inspector.get_columns("parser_profiles")}
if "kind" not in columns:
statements.append(
"ALTER TABLE parser_profiles ADD COLUMN kind VARCHAR(50) NOT NULL DEFAULT 'heuristic'"
)
if "llm_profile" not in columns:
statements.append("ALTER TABLE parser_profiles ADD COLUMN llm_profile JSON")
if not statements:
return
@@ -0,0 +1,310 @@
"""Parser builder: one-shot DeepSeek profile generation + static preview."""
from __future__ import annotations
import json
import logging
import os
import re
from typing import Any
import httpx
from contracts.heuristic_profile import (
HeuristicProfile,
TARGET_FIELDS,
apply_profile,
match_profile,
target_field_specs,
)
logger = logging.getLogger("ca.parser_builder")
GENERATE_SYSTEM = (
"You design a static heuristic parser profile for structured Telegram posts. "
"Output valid JSON only, no markdown, no Python code. "
"The profile is applied with regex/line/marker rules at runtime — never with an LLM. "
"Every rule must actually extract a non-empty value from the given sample when applied."
)
def deepseek_enabled() -> bool:
return bool(os.getenv("DEEPSEEK_API_KEY", "").strip())
def deepseek_settings() -> dict[str, str]:
return {
"api_key": os.getenv("DEEPSEEK_API_KEY", "").strip(),
"base_url": os.getenv("DEEPSEEK_BASE_URL", "https://api.deepseek.com").rstrip("/"),
"model": os.getenv("DEEPSEEK_MODEL", "deepseek-chat"),
}
def get_target_fields() -> list[dict[str, str]]:
return target_field_specs()
def preview_with_profile(sample_post: str, profile: dict[str, Any] | HeuristicProfile) -> dict[str, str]:
return apply_profile(sample_post, profile)
def match_preview(
sample_post: str,
profile: dict[str, Any] | HeuristicProfile,
) -> tuple[dict[str, str], bool, list[str]]:
matched, fields, missing = match_profile(sample_post, profile)
return fields, matched, missing
def empty_preview_fields(preview: dict[str, str]) -> list[str]:
return [name for name, value in preview.items() if not (value or "").strip()]
def _extract_json_object(content: str) -> dict[str, Any]:
content = (content or "").strip()
if content.startswith("```"):
content = re.sub(r"^```(?:json)?\s*", "", content)
content = re.sub(r"\s*```$", "", content)
try:
parsed = json.loads(content)
except json.JSONDecodeError:
match = re.search(r"\{[\s\S]*\}", content)
if not match:
raise
parsed = json.loads(match.group(0))
if not isinstance(parsed, dict):
raise ValueError("LLM response must be a JSON object")
return parsed
def _base_user_prompt(sample: str) -> str:
fields_help = "\n".join(
f"- {spec['name']}: {spec['description']}" for spec in target_field_specs()
)
return (
"Build a HeuristicProfile JSON for this sample Telegram post.\n\n"
"Schema:\n"
'{"version": 1, "notes": "...", "fields": {'
'"<field>": {"strategy": "regex|line|after_marker|between|full_text|literal", '
'"pattern": "...", "group": 1, "line_index": 0, "marker": "...", '
'"end_marker": "...", "value": "...", "flags": "im", "strip": true}'
"}}\n\n"
"Rules:\n"
"- Only include fields you can extract reliably from the sample.\n"
f"- Allowed field names: {', '.join(TARGET_FIELDS)}.\n"
"- Prefer regex / line / after_marker / between over literal.\n"
"- Use literal only for constant topic/region tags.\n"
"- title: usually the first non-empty line (strategy line, line_index 0).\n"
"- description: the multi-line body AFTER the title and BEFORE coords/hashtags. "
"Prefer strategy between or regex with flags including \"s\" (DOTALL). "
"Do NOT leave description empty if the sample has a body paragraph.\n"
"- locality: settlement name from the title line when present.\n"
"- coords should capture 'lat, lon' when present.\n"
"- event_date should capture DD.MM.YYYY / DD.MM.YY or similar.\n"
"- topic/region: hashtags like #ru → topic \"ru\"; avoid stuffing unrelated tags into region.\n"
"- Do not invent Python code; only declarative rules.\n"
"- Mentally apply each rule to the sample: every included field must yield a non-empty string.\n\n"
f"Target fields:\n{fields_help}\n\n"
f"Sample post:\n{sample[:12000]}"
)
def _parse_profile_response(content: str) -> HeuristicProfile:
raw = _extract_json_object(content)
if "fields" not in raw and isinstance(raw.get("profile"), dict):
raw = raw["profile"]
if "version" not in raw:
raw["version"] = 1
return HeuristicProfile.model_validate(raw)
async def generate_profile(
sample_post: str,
*,
hint: str | None = None,
current_profile: dict[str, Any] | HeuristicProfile | None = None,
) -> HeuristicProfile:
settings = deepseek_settings()
if not settings["api_key"]:
raise RuntimeError(
"DEEPSEEK_API_KEY is not set on ca-api. Add it to .env for parser generation."
)
sample = (sample_post or "").strip()
if not sample:
raise ValueError("sample_post is required")
hint_text = (hint or "").strip()
current: HeuristicProfile | None = None
if current_profile is not None:
current = (
current_profile
if isinstance(current_profile, HeuristicProfile)
else HeuristicProfile.model_validate(current_profile)
)
if current is not None or hint_text:
preview = preview_with_profile(sample, current) if current is not None else {}
empty = empty_preview_fields(preview) if preview else []
parts = [
"Revise the HeuristicProfile JSON for this sample Telegram post.",
"Keep strategies that already work; fix only what the manager asks for "
"and any fields that preview as empty.",
"Output the full updated profile JSON (same schema), not a patch.",
"",
_base_user_prompt(sample),
]
if current is not None:
parts.extend(
[
"",
"Current profile JSON:",
json.dumps(current.model_dump(), ensure_ascii=False, indent=2)[:12000],
]
)
if empty:
parts.append(
"Fields currently empty in preview (must become non-empty if present in sample): "
+ ", ".join(empty)
)
else:
parts.append("Current preview fields: " + json.dumps(preview, ensure_ascii=False))
if hint_text:
parts.extend(["", "Manager hint (follow this):", hint_text[:4000]])
user_prompt = "\n".join(parts)
else:
user_prompt = _base_user_prompt(sample)
payload = {
"model": settings["model"],
"messages": [
{"role": "system", "content": GENERATE_SYSTEM},
{"role": "user", "content": user_prompt},
],
"temperature": 0.1,
"response_format": {"type": "json_object"},
}
url = f"{settings['base_url']}/chat/completions"
async with httpx.AsyncClient(timeout=90.0) as client:
response = await client.post(
url,
headers={
"Authorization": f"Bearer {settings['api_key']}",
"Content-Type": "application/json",
},
json=payload,
)
response.raise_for_status()
data = response.json()
content = data["choices"][0]["message"]["content"]
return _parse_profile_response(content)
async def preview_llm_extract(
sample_post: str,
llm_profile: dict[str, Any],
) -> tuple[list[dict[str, str]], dict[str, str], bool, list[str], bool, int]:
"""One-shot DeepSeek extract for CA preview (does not persist).
Returns (events, fields, matched, missing_required, is_event, matched_count).
``fields`` is the first event (or empty) for legacy UI compatibility.
"""
from contracts.llm_profile import DEFAULT_INSTRUCTION, LlmProfile, match_llm_required
settings = deepseek_settings()
if not settings["api_key"]:
raise RuntimeError(
"DEEPSEEK_API_KEY is not set on ca-api. Add it to .env for LLM preview."
)
sample = (sample_post or "").strip()
if not sample:
raise ValueError("sample_post is required")
profile = LlmProfile.model_validate(llm_profile)
schema = profile.extract_schema
instr = profile.instruction or DEFAULT_INSTRUCTION
schema_lines = "\n".join(f"- {k}: {v}" for k, v in schema.items())
multi = bool(profile.multi_event)
if multi:
user_prompt = (
f"{instr}\n\n"
"If the text describes multiple distinct events (different places, "
"coords, or dates), return one object per event in \"events\".\n"
f"Fields per event:\n{schema_lines}\n\n"
'Return JSON: {"is_event": true|false, "events": [{<field>: <string>}, ...]}\n'
"If there is no event, return is_event=false and events=[].\n\n"
f"Text:\n{sample[:12000]}"
)
else:
user_prompt = (
f"{instr}\n\n"
f"Fields to extract:\n{schema_lines}\n\n"
'Return JSON: {"is_event": true|false, "fields": {<field>: <string>}}\n\n'
f"Text:\n{sample[:12000]}"
)
payload = {
"model": settings["model"],
"messages": [
{
"role": "system",
"content": (
"You extract structured event data for a geoint map. "
"Output valid JSON only, no markdown."
),
},
{"role": "user", "content": user_prompt},
],
"temperature": 0.1,
"response_format": {"type": "json_object"},
}
url = f"{settings['base_url']}/chat/completions"
async with httpx.AsyncClient(timeout=90.0) as client:
response = await client.post(
url,
headers={
"Authorization": f"Bearer {settings['api_key']}",
"Content-Type": "application/json",
},
json=payload,
)
response.raise_for_status()
data = response.json()
content = data["choices"][0]["message"]["content"]
parsed = _extract_json_object(content)
is_event = bool(parsed.get("is_event", True))
empty_fields = {key: "" for key in schema}
if not is_event:
return [], empty_fields, False, [], False, 0
events: list[dict[str, str]] = []
if multi and isinstance(parsed.get("events"), list):
for item in parsed["events"]:
if isinstance(item, dict):
events.append({key: str(item.get(key) or "").strip() for key in schema})
else:
fields_raw = parsed.get("fields") if isinstance(parsed.get("fields"), dict) else parsed
if not isinstance(fields_raw, dict):
fields_raw = {}
events.append({key: str(fields_raw.get(key) or "").strip() for key in schema})
if not events:
return [], empty_fields, False, [], False, 0
matched_events = [ev for ev in events if match_llm_required(ev, list(profile.required_fields))]
fields = events[0]
missing = [
name
for name in profile.required_fields
if not str(fields.get(name) or "").strip()
]
matched = match_llm_required(fields, list(profile.required_fields))
return events, fields, matched, missing, True, len(matched_events)
@@ -5,7 +5,8 @@ from datetime import datetime, timezone
from ..database import SessionLocal
from ..models import ParseJob
from .jobs import enqueue_job
from .job_stale import STALE_JOB_SECONDS, is_stale_job
from .jobs import enqueue_parse_job
logger = logging.getLogger(__name__)
@@ -13,10 +14,47 @@ TICK_SECONDS = 30
RECURRING_STATUSES = ("completed", "failed")
def recover_stale_jobs(db, now: datetime) -> None:
"""Mark abandoned running/queued jobs as failed; re-queue if still active."""
stuck = (
db.query(ParseJob)
.filter(ParseJob.status.in_(("running", "queued")))
.all()
)
for job in stuck:
if not is_stale_job(job, now):
continue
prev = job.status
job.status = "failed"
job.last_error = (
f"Stale {prev} recovered after {STALE_JOB_SECONDS}s "
"(worker likely restarted)"
)
job.last_run_at = now
db.commit()
logger.warning("Recovered stale job %s (was %s)", job.id, prev)
if not job.is_active or job.interval_seconds <= 0:
continue
job.status = "queued"
job.last_error = None
db.commit()
try:
enqueue_parse_job(db, job)
logger.info("Re-queued recovered job %s", job.id)
except ValueError as exc:
job.status = "failed"
job.last_error = str(exc)
db.commit()
logger.warning("Skip re-queue recovered job %s: %s", job.id, exc)
def run_scheduler_tick() -> None:
db = SessionLocal()
try:
now = datetime.now(timezone.utc)
recover_stale_jobs(db, now)
jobs = (
db.query(ParseJob)
.filter(
@@ -36,7 +74,14 @@ def run_scheduler_tick() -> None:
job.status = "queued"
job.last_error = None
db.commit()
enqueue_job(job.id, job.source_type, job.source_config)
try:
enqueue_parse_job(db, job)
except ValueError as exc:
job.status = "failed"
job.last_error = str(exc)
db.commit()
logger.warning("Skip re-queue job %s: %s", job.id, exc)
continue
logger.info("Re-queued recurring job %s (interval %ss)", job.id, job.interval_seconds)
except Exception:
logger.exception("Scheduler tick failed")
@@ -54,5 +99,9 @@ def start_scheduler() -> threading.Event:
stop_event = threading.Event()
thread = threading.Thread(target=_scheduler_loop, args=(stop_event,), daemon=True)
thread.start()
logger.info("Parse job scheduler started (tick every %ss)", TICK_SECONDS)
logger.info(
"Parse job scheduler started (tick every %ss, stale after %ss)",
TICK_SECONDS,
STALE_JOB_SECONDS,
)
return stop_event
+137
View File
@@ -0,0 +1,137 @@
"""VPN settings persistence helpers."""
from __future__ import annotations
from contracts.queues import known_source_types
from contracts.vpn import VpnSettings
from sqlalchemy.orm import Session
from ..models import VpnSettingsRow
from ..schemas import VpnSettingsAdminRead, VpnSettingsInternalRead, VpnSettingsUpdate
SINGLETON_ID = 1
def _hint_secret(value: str | None) -> str | None:
if not value:
return None
if len(value) <= 4:
return "****"
return f"…{value[-4:]}"
def ensure_vpn_settings(db: Session) -> VpnSettingsRow:
row = db.query(VpnSettingsRow).filter(VpnSettingsRow.id == SINGLETON_ID).first()
if row is not None:
return row
row = VpnSettingsRow(
id=SINGLETON_ID,
enabled=False,
mode="subscription",
subscription_interval_seconds=3600,
proxied_source_types=[],
)
db.add(row)
db.commit()
db.refresh(row)
return row
def row_to_contract(row: VpnSettingsRow) -> VpnSettings:
return VpnSettings.model_validate(
{
"enabled": row.enabled,
"mode": row.mode,
"subscription_url": row.subscription_url,
"subscription_interval_seconds": row.subscription_interval_seconds,
"host": row.host,
"port": row.port,
"username": row.username,
"password": row.password,
"proxied_source_types": row.proxied_source_types or [],
}
)
def to_admin_read(row: VpnSettingsRow) -> VpnSettingsAdminRead:
return VpnSettingsAdminRead(
enabled=row.enabled,
mode=row.mode, # type: ignore[arg-type]
subscription_url_set=bool(row.subscription_url),
subscription_url_hint=_hint_secret(row.subscription_url),
subscription_interval_seconds=row.subscription_interval_seconds,
host=row.host,
port=row.port,
username=row.username,
password_set=bool(row.password),
proxied_source_types=list(row.proxied_source_types or []),
available_source_types=known_source_types(),
)
def to_internal_read(row: VpnSettingsRow) -> VpnSettingsInternalRead:
return VpnSettingsInternalRead(
enabled=row.enabled,
mode=row.mode, # type: ignore[arg-type]
subscription_url=row.subscription_url,
subscription_interval_seconds=row.subscription_interval_seconds,
host=row.host,
port=row.port,
username=row.username,
password=row.password,
proxied_source_types=list(row.proxied_source_types or []),
)
def apply_vpn_update(db: Session, payload: VpnSettingsUpdate) -> VpnSettingsRow:
row = ensure_vpn_settings(db)
data = {
"enabled": row.enabled if payload.enabled is None else payload.enabled,
"mode": row.mode if payload.mode is None else payload.mode,
"subscription_url": row.subscription_url,
"subscription_interval_seconds": (
row.subscription_interval_seconds
if payload.subscription_interval_seconds is None
else payload.subscription_interval_seconds
),
"host": row.host if payload.host is None else (payload.host.strip() or None),
"port": row.port if payload.port is None else payload.port,
"username": (
row.username
if payload.username is None
else (payload.username.strip() or None)
),
"password": row.password,
"proxied_source_types": (
list(row.proxied_source_types or [])
if payload.proxied_source_types is None
else payload.proxied_source_types
),
}
if payload.clear_subscription_url:
data["subscription_url"] = None
elif payload.subscription_url is not None:
stripped = payload.subscription_url.strip()
data["subscription_url"] = stripped or None
if payload.clear_password:
data["password"] = None
elif payload.password is not None:
stripped = payload.password.strip()
data["password"] = stripped or None
validated = VpnSettings.model_validate(data)
row.enabled = validated.enabled
row.mode = validated.mode
row.subscription_url = validated.subscription_url
row.subscription_interval_seconds = validated.subscription_interval_seconds
row.host = validated.host
row.port = validated.port
row.username = validated.username
row.password = validated.password
row.proxied_source_types = validated.proxied_source_types
db.commit()
db.refresh(row)
return row
+2
View File
@@ -5,3 +5,5 @@ pydantic==2.10.3
python-multipart==0.0.20
psycopg2-binary==2.9.10
redis==5.2.1
httpx==0.28.1
PyJWT==2.10.1
@@ -4,14 +4,18 @@ import type {
Consumer,
ConsumerCreate,
ConsumerUpdate,
EventCreate,
EventFilters,
EventListResponse,
EventRecord,
EventUpdate,
ParseJob,
ParseJobCreate,
ParseJobUpdate,
TimelinePoint,
TopItem,
VpnSettings,
VpnSettingsUpdate,
} from "../types/admin";
function buildQuery(params: Record<string, string | number | undefined>): string {
@@ -69,6 +73,26 @@ export function fetchEvents(filters: EventFilters = {}): Promise<EventListRespon
);
}
export function createEvent(payload: EventCreate): Promise<EventRecord> {
return request<EventRecord>(
"/events",
{ method: "POST", body: JSON.stringify(payload) },
ADMIN_BASE,
);
}
export function updateEvent(id: number, payload: EventUpdate): Promise<EventRecord> {
return request<EventRecord>(
`/events/${id}`,
{ method: "PATCH", body: JSON.stringify(payload) },
ADMIN_BASE,
);
}
export function deleteEvent(id: number): Promise<void> {
return request<void>(`/events/${id}`, { method: "DELETE" }, ADMIN_BASE);
}
export function fetchConsumers(): Promise<Consumer[]> {
return request<Consumer[]>("/consumers", undefined, ADMIN_BASE);
}
@@ -135,3 +159,15 @@ export async function testDistribution(apiKey: string, limit = 5): Promise<Event
}
return response.json() as Promise<EventRecord[]>;
}
export function fetchVpnSettings(): Promise<VpnSettings> {
return request<VpnSettings>("/vpn", undefined, ADMIN_BASE);
}
export function updateVpnSettings(payload: VpnSettingsUpdate): Promise<VpnSettings> {
return request<VpnSettings>(
"/vpn",
{ method: "PUT", body: JSON.stringify(payload) },
ADMIN_BASE,
);
}
@@ -1,3 +1,5 @@
import { clearToken, getToken } from "../auth";
const API_BASE = "/api";
const ADMIN_BASE = "/admin";
@@ -13,11 +15,24 @@ export async function request<T>(
headers.set("Content-Type", "application/json");
}
const token = getToken();
if (token && !headers.has("Authorization")) {
headers.set("Authorization", `Bearer ${token}`);
}
const response = await fetch(`${base}${url}`, {
...options,
headers,
});
if (response.status === 401 && token) {
clearToken();
if (typeof window !== "undefined" && !window.location.pathname.startsWith("/login")) {
const next = `${window.location.pathname}${window.location.search}`;
window.location.assign(`/login?next=${encodeURIComponent(next)}`);
}
}
if (!response.ok) {
const message = await response.text();
throw new Error(message || `Ошибка запроса: ${response.status}`);
@@ -0,0 +1,133 @@
import { ADMIN_BASE, request } from "./client";
import type {
HeuristicProfile,
LlmProfile,
ParseChannel,
ParseChannelCreate,
ParseChannelUpdate,
ParserProfile,
ParserProfileCreate,
ParserProfileUpdate,
TargetField,
} from "../types/admin";
export type { HeuristicProfile, LlmProfile, TargetField } from "../types/admin";
export function fetchChannels(): Promise<ParseChannel[]> {
return request<ParseChannel[]>("/parse-channels", undefined, ADMIN_BASE);
}
export function createChannel(payload: ParseChannelCreate): Promise<ParseChannel> {
return request<ParseChannel>(
"/parse-channels",
{ method: "POST", body: JSON.stringify(payload) },
ADMIN_BASE,
);
}
export function updateChannel(id: number, payload: ParseChannelUpdate): Promise<ParseChannel> {
return request<ParseChannel>(
`/parse-channels/${id}`,
{ method: "PATCH", body: JSON.stringify(payload) },
ADMIN_BASE,
);
}
export function deleteChannel(id: number): Promise<void> {
return request<void>(`/parse-channels/${id}`, { method: "DELETE" }, ADMIN_BASE);
}
export function fetchProfiles(): Promise<ParserProfile[]> {
return request<ParserProfile[]>("/parser-profiles", undefined, ADMIN_BASE);
}
export function createProfile(payload: ParserProfileCreate): Promise<ParserProfile> {
return request<ParserProfile>(
"/parser-profiles",
{ method: "POST", body: JSON.stringify(payload) },
ADMIN_BASE,
);
}
export function updateProfile(id: number, payload: ParserProfileUpdate): Promise<ParserProfile> {
return request<ParserProfile>(
`/parser-profiles/${id}`,
{ method: "PATCH", body: JSON.stringify(payload) },
ADMIN_BASE,
);
}
export function deleteProfile(id: number): Promise<void> {
return request<void>(`/parser-profiles/${id}`, { method: "DELETE" }, ADMIN_BASE);
}
export function fetchTargetFields(): Promise<{ fields: TargetField[] }> {
return request<{ fields: TargetField[] }>("/parser-profiles/target-fields", undefined, ADMIN_BASE);
}
export function fetchLlmDefaults(): Promise<LlmProfile> {
return request<LlmProfile>("/parser-profiles/llm-defaults", undefined, ADMIN_BASE);
}
export function generateParserProfile(payload: {
sample_post: string;
hint?: string;
current_profile?: HeuristicProfile | Record<string, unknown> | null;
}): Promise<{
profile: HeuristicProfile;
preview: Record<string, string>;
empty_fields: string[];
}> {
return request<{
profile: HeuristicProfile;
preview: Record<string, string>;
empty_fields: string[];
}>(
"/parser-profiles/generate",
{
method: "POST",
body: JSON.stringify({
sample_post: payload.sample_post,
hint: payload.hint || undefined,
current_profile: payload.current_profile || undefined,
}),
},
ADMIN_BASE,
);
}
export function previewParserProfile(
sample_post: string,
heuristic_profile: HeuristicProfile | Record<string, unknown>,
): Promise<{ fields: Record<string, string>; matched: boolean; missing_required: string[] }> {
return request<{ fields: Record<string, string>; matched: boolean; missing_required: string[] }>(
"/parser-profiles/preview",
{ method: "POST", body: JSON.stringify({ sample_post, heuristic_profile }) },
ADMIN_BASE,
);
}
export function previewLlmProfile(
sample_post: string,
llm_profile: LlmProfile | Record<string, unknown>,
): Promise<{
fields: Record<string, string>;
events: Record<string, string>[];
matched: boolean;
missing_required: string[];
is_event: boolean;
matched_count: number;
}> {
return request<{
fields: Record<string, string>;
events: Record<string, string>[];
matched: boolean;
missing_required: string[];
is_event: boolean;
matched_count: number;
}>(
"/parser-profiles/preview-llm",
{ method: "POST", body: JSON.stringify({ sample_post, llm_profile }) },
ADMIN_BASE,
);
}
+40
View File
@@ -0,0 +1,40 @@
import { computed, ref } from "vue";
const TOKEN_KEY = "mapmil_admin_token";
const ADMIN_BASE = "/admin";
const token = ref<string | null>(localStorage.getItem(TOKEN_KEY));
export function getToken(): string | null {
return token.value;
}
export function setToken(value: string | null): void {
token.value = value;
if (value) localStorage.setItem(TOKEN_KEY, value);
else localStorage.removeItem(TOKEN_KEY);
}
export function clearToken(): void {
setToken(null);
}
export const isAuthenticated = computed(() => Boolean(token.value));
export async function login(username: string, password: string): Promise<void> {
const response = await fetch(`${ADMIN_BASE}/auth/login`, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ username, password }),
});
if (!response.ok) {
const message = await response.text();
throw new Error(message || `Ошибка входа: ${response.status}`);
}
const data = (await response.json()) as { access_token: string };
setToken(data.access_token);
}
export function logout(): void {
clearToken();
}
@@ -4,11 +4,15 @@ import { onMounted, onUnmounted, watch } from "vue";
import { useLeafletMap } from "../composables/useLeafletMap";
import type { MapObjectWithEvent } from "../types/map";
const props = defineProps<{
objects: MapObjectWithEvent[];
selectedId: number | null;
openPopupId: number | null;
}>();
const props = withDefaults(
defineProps<{
objects: MapObjectWithEvent[];
selectedId: number | null;
openPopupId: number | null;
editable?: boolean;
}>(),
{ editable: false },
);
const emit = defineEmits<{
select: [object: MapObjectWithEvent];
@@ -53,7 +57,7 @@ function buildPopupHtml(obj: MapObjectWithEvent): string {
}
function isDraggable(obj: MapObjectWithEvent): boolean {
return obj.id === props.selectedId && !obj.event_id;
return props.editable && obj.id === props.selectedId && !obj.event_id;
}
function syncMarkers() {
@@ -170,7 +174,7 @@ onUnmounted(() => {
});
watch(
() => [props.objects, props.selectedId] as const,
() => [props.objects, props.selectedId, props.editable] as const,
() => {
syncMarkers();
panToSelected(false);
@@ -9,9 +9,13 @@ import {
import { MEDIA_ACCEPT, OBJECT_TYPE_LABELS } from "../types/object";
import type { MapObjectWithEvent } from "../types/map";
const props = defineProps<{
object: MapObjectWithEvent | null;
}>();
const props = withDefaults(
defineProps<{
object: MapObjectWithEvent | null;
canEdit?: boolean;
}>(),
{ canEdit: false },
);
const emit = defineEmits<{
openEvents: [];
@@ -151,7 +155,7 @@ watch(() => props.object?.id, () => loadMedia(), { immediate: true });
<section v-if="!isIngested" class="media-section">
<div class="media-header">
<h4>Медиа</h4>
<label class="upload-btn">
<label v-if="canEdit" class="upload-btn">
{{ uploading ? "Загрузка..." : "Добавить" }}
<input
type="file"
@@ -185,7 +189,14 @@ watch(() => props.object?.id, () => loadMedia(), { immediate: true });
<span class="name" :title="item.original_name">{{ item.original_name }}</span>
<span class="size">{{ formatFileSize(item.size) }}</span>
</div>
<button type="button" class="delete-btn" @click="handleDelete(item.id)">Удалить</button>
<button
v-if="canEdit"
type="button"
class="delete-btn"
@click="handleDelete(item.id)"
>
Удалить
</button>
</li>
</ul>
</section>
@@ -0,0 +1,41 @@
import { computed, ref } from "vue";
import { useRouter } from "vue-router";
import { clearToken, getToken, isAuthenticated, login as doLogin, logout as doLogout } from "../auth";
export function useAuth() {
const router = useRouter();
const busy = ref(false);
const error = ref("");
const authenticated = computed(() => isAuthenticated.value);
async function login(username: string, password: string): Promise<boolean> {
busy.value = true;
error.value = "";
try {
await doLogin(username, password);
return true;
} catch (err) {
clearToken();
error.value = err instanceof Error ? err.message : "Не удалось войти";
return false;
} finally {
busy.value = false;
}
}
function logout(): void {
doLogout();
void router.push({ name: "map" });
}
return {
authenticated,
busy,
error,
getToken,
login,
logout,
};
}
@@ -2,23 +2,34 @@
import { computed } from "vue";
import { useRoute } from "vue-router";
const route = useRoute();
import { useAuth } from "../composables/useAuth";
const navItems = [
{ to: "/", label: "Карта", exact: true },
const route = useRoute();
const { authenticated, logout } = useAuth();
const publicNavItems = [{ to: "/", label: "Карта", exact: true }];
const adminNavItems = [
{ to: "/channels", label: "Каналы" },
{ to: "/parser-profiles", label: "Профили" },
{ to: "/parsers", label: "Парсеры" },
{ to: "/vpn", label: "VPN" },
{ to: "/events", label: "События" },
{ to: "/analytics", label: "Аналитика" },
{ to: "/consumers", label: "ПИ" },
];
const navItems = computed(() =>
authenticated.value ? [...publicNavItems, ...adminNavItems] : publicNavItems,
);
function isActive(path: string, exact = false): boolean {
if (exact) return route.path === path;
return route.path.startsWith(path);
}
const pageTitle = computed(() => {
const item = navItems.find((n) => isActive(n.to, n.exact));
const item = navItems.value.find((n) => isActive(n.to, n.exact));
return item?.label ?? "MapMil";
});
</script>
@@ -40,6 +51,12 @@ const pageTitle = computed(() => {
>
{{ item.label }}
</router-link>
<router-link v-if="!authenticated" to="/login" class="nav-link nav-auth">
Вход
</router-link>
<button v-else type="button" class="nav-link nav-auth nav-button" @click="logout">
Выход
</button>
</nav>
</header>
<main class="admin-main">
@@ -85,6 +102,7 @@ const pageTitle = computed(() => {
.nav {
display: flex;
align-items: center;
gap: 0.25rem;
}
@@ -107,6 +125,19 @@ const pageTitle = computed(() => {
font-weight: 500;
}
.nav-auth {
margin-left: 0.5rem;
color: #1565c0;
font-weight: 500;
}
.nav-button {
border: none;
background: transparent;
cursor: pointer;
font: inherit;
}
.admin-main {
flex: 1;
min-height: 0;
+63 -5
View File
@@ -1,27 +1,85 @@
import { createRouter, createWebHistory } from "vue-router";
import { getToken } from "../auth";
import AppLayout from "../layouts/AppLayout.vue";
import AnalyticsView from "../views/AnalyticsView.vue";
import ChannelsView from "../views/ChannelsView.vue";
import ConsumersView from "../views/ConsumersView.vue";
import EventsView from "../views/EventsView.vue";
import LoginView from "../views/LoginView.vue";
import MapViewPage from "../views/MapViewPage.vue";
import ParserProfilesView from "../views/ParserProfilesView.vue";
import ParsersView from "../views/ParsersView.vue";
import VpnView from "../views/VpnView.vue";
const router = createRouter({
history: createWebHistory(),
routes: [
{
path: "/login",
name: "login",
component: LoginView,
meta: { public: true },
},
{
path: "/",
component: AppLayout,
children: [
{ path: "", name: "map", component: MapViewPage },
{ path: "parsers", name: "parsers", component: ParsersView },
{ path: "events", name: "events", component: EventsView },
{ path: "analytics", name: "analytics", component: AnalyticsView },
{ path: "consumers", name: "consumers", component: ConsumersView },
{ path: "", name: "map", component: MapViewPage, meta: { public: true } },
{
path: "channels",
name: "channels",
component: ChannelsView,
meta: { requiresAuth: true },
},
{
path: "parser-profiles",
name: "parser-profiles",
component: ParserProfilesView,
meta: { requiresAuth: true },
},
{
path: "parsers",
name: "parsers",
component: ParsersView,
meta: { requiresAuth: true },
},
{
path: "vpn",
name: "vpn",
component: VpnView,
meta: { requiresAuth: true },
},
{
path: "events",
name: "events",
component: EventsView,
meta: { requiresAuth: true },
},
{
path: "analytics",
name: "analytics",
component: AnalyticsView,
meta: { requiresAuth: true },
},
{
path: "consumers",
name: "consumers",
component: ConsumersView,
meta: { requiresAuth: true },
},
],
},
],
});
router.beforeEach((to) => {
if (!to.meta.requiresAuth) return true;
if (getToken()) return true;
return {
name: "login",
query: { next: to.fullPath },
};
});
export default router;
+9
View File
@@ -125,6 +125,15 @@ textarea {
font-size: 0.8rem;
}
.btn-danger {
color: #b71c1c;
border-color: #ef9a9a;
}
.btn-danger:hover {
background: #ffebee;
}
.link-btn {
background: none;
border: none;
+147 -2
View File
@@ -2,6 +2,8 @@ export interface ParseJob {
id: number;
source_type: string;
source_config: Record<string, unknown>;
profile_id?: number | null;
channel_id?: number | null;
schedule: string | null;
interval_seconds: number;
is_active: boolean;
@@ -9,11 +11,17 @@ export interface ParseJob {
last_run_at: string | null;
last_error: string | null;
created_at: string;
profile_name?: string | null;
channel_name?: string | null;
channel_handle?: string | null;
}
export interface ParseJobCreate {
source_type: string;
source_config: Record<string, unknown>;
source_type?: string;
source_config?: Record<string, unknown>;
profile_id?: number;
channel_id?: number;
limit?: number;
schedule?: string | null;
interval_seconds?: number;
is_active?: boolean;
@@ -21,10 +29,85 @@ export interface ParseJobCreate {
export interface ParseJobUpdate {
source_config?: Record<string, unknown>;
profile_id?: number;
channel_id?: number;
limit?: number;
interval_seconds?: number;
is_active?: boolean;
}
export type TargetField = {
name: string;
type: string;
description: string;
};
export type HeuristicProfile = {
version: number;
fields: Record<string, Record<string, unknown>>;
required_fields?: string[];
notes?: string;
};
export type LlmProfile = {
instruction?: string | null;
extract_schema: Record<string, string>;
required_fields?: string[];
multi_event?: boolean;
};
export interface ParserProfile {
id: number;
name: string;
kind: "heuristic" | "llm";
sample_post: string;
heuristic_profile: Record<string, unknown> | null;
llm_profile: LlmProfile | Record<string, unknown> | null;
status: string;
created_at: string;
}
export interface ParserProfileCreate {
name: string;
kind?: "heuristic" | "llm";
sample_post?: string;
heuristic_profile?: Record<string, unknown> | null;
llm_profile?: LlmProfile | Record<string, unknown> | null;
status?: string;
}
export interface ParserProfileUpdate {
name?: string;
kind?: "heuristic" | "llm";
sample_post?: string;
heuristic_profile?: Record<string, unknown> | null;
llm_profile?: LlmProfile | Record<string, unknown> | null;
status?: string;
}
export interface ParseChannel {
id: number;
name: string;
source_type: string;
channel: string;
is_active: boolean;
created_at: string;
}
export interface ParseChannelCreate {
name: string;
source_type?: string;
channel: string;
is_active?: boolean;
}
export interface ParseChannelUpdate {
name?: string;
source_type?: string;
channel?: string;
is_active?: boolean;
}
export interface EventRecord {
id: number;
source_type: string;
@@ -43,6 +126,38 @@ export interface EventRecord {
metadata: Record<string, unknown> | null;
}
export interface EventCreate {
source_type?: string;
source_url: string;
raw_text?: string;
title?: string;
description?: string;
locality?: string;
latitude?: number | null;
longitude?: number | null;
event_date?: string | null;
region?: string | null;
topic?: string | null;
tags?: string[] | null;
metadata?: Record<string, unknown> | null;
}
export interface EventUpdate {
source_type?: string;
source_url?: string;
raw_text?: string;
title?: string;
description?: string;
locality?: string;
latitude?: number | null;
longitude?: number | null;
event_date?: string | null;
region?: string | null;
topic?: string | null;
tags?: string[] | null;
metadata?: Record<string, unknown> | null;
}
export interface EventListResponse {
items: EventRecord[];
total: number;
@@ -105,3 +220,33 @@ export interface TopItem {
name: string;
count: number;
}
export type VpnMode = "subscription" | "socks5" | "http";
export interface VpnSettings {
enabled: boolean;
mode: VpnMode;
subscription_url_set: boolean;
subscription_url_hint: string | null;
subscription_interval_seconds: number;
host: string | null;
port: number | null;
username: string | null;
password_set: boolean;
proxied_source_types: string[];
available_source_types: string[];
}
export interface VpnSettingsUpdate {
enabled?: boolean;
mode?: VpnMode;
subscription_url?: string | null;
clear_subscription_url?: boolean;
subscription_interval_seconds?: number;
host?: string | null;
port?: number | null;
username?: string | null;
password?: string | null;
clear_password?: boolean;
proxied_source_types?: string[];
}
@@ -0,0 +1,273 @@
<script setup lang="ts">
import { onMounted, ref } from "vue";
import {
createChannel,
deleteChannel,
fetchChannels,
updateChannel,
} from "../api/entities";
import type { ParseChannel } from "../types/admin";
const channels = ref<ParseChannel[]>([]);
const loading = ref(true);
const error = ref("");
const submitting = ref(false);
const form = ref({
name: "",
channel: "",
is_active: true,
});
const editModal = ref<{
open: boolean;
channel: ParseChannel | null;
name: string;
handle: string;
is_active: boolean;
}>({
open: false,
channel: null,
name: "",
handle: "",
is_active: true,
});
async function loadChannels() {
loading.value = true;
error.value = "";
try {
channels.value = await fetchChannels();
} catch (err) {
error.value = err instanceof Error ? err.message : "Не удалось загрузить каналы";
} finally {
loading.value = false;
}
}
async function handleCreate() {
if (!form.value.name.trim() || !form.value.channel.trim()) {
error.value = "Укажите название и handle канала";
return;
}
submitting.value = true;
error.value = "";
try {
await createChannel({
name: form.value.name.trim(),
channel: form.value.channel.trim(),
source_type: "telegram",
is_active: form.value.is_active,
});
form.value = { name: "", channel: "", is_active: true };
await loadChannels();
} catch (err) {
error.value = err instanceof Error ? err.message : "Не удалось создать канал";
} finally {
submitting.value = false;
}
}
function openEdit(ch: ParseChannel) {
editModal.value = {
open: true,
channel: ch,
name: ch.name,
handle: ch.channel,
is_active: ch.is_active,
};
}
function closeEdit() {
editModal.value.open = false;
editModal.value.channel = null;
}
async function handleSaveEdit() {
if (!editModal.value.channel) return;
if (!editModal.value.name.trim() || !editModal.value.handle.trim()) {
error.value = "Укажите название и handle";
return;
}
submitting.value = true;
error.value = "";
try {
await updateChannel(editModal.value.channel.id, {
name: editModal.value.name.trim(),
channel: editModal.value.handle.trim(),
is_active: editModal.value.is_active,
});
closeEdit();
await loadChannels();
} catch (err) {
error.value = err instanceof Error ? err.message : "Не удалось сохранить";
} finally {
submitting.value = false;
}
}
async function handleDelete(ch: ParseChannel) {
if (!window.confirm(`Удалить канал «${ch.name}» (@${ch.channel})?`)) return;
error.value = "";
try {
await deleteChannel(ch.id);
await loadChannels();
} catch (err) {
error.value = err instanceof Error ? err.message : "Не удалось удалить";
}
}
function formatDate(value: string): string {
return new Date(value).toLocaleString("ru-RU");
}
onMounted(loadChannels);
</script>
<template>
<div class="page">
<h2 class="page-heading">Каналы</h2>
<p class="intro">
Каналы для парсинга (Telegram). Связка канал + профиль создаётся на странице
<router-link to="/parsers">Парсеры</router-link>.
</p>
<section class="card">
<h3>Новый канал</h3>
<form class="admin-form" @submit.prevent="handleCreate">
<div class="form-row">
<label>
Название
<input v-model="form.name" type="text" placeholder="Сводки фронта" required />
</label>
<label>
Handle / id
<input v-model="form.channel" type="text" placeholder="example_channel" required />
</label>
<label class="checkbox-label">
<input v-model="form.is_active" type="checkbox" />
Активен
</label>
</div>
<div class="form-actions">
<button type="submit" class="btn btn-primary" :disabled="submitting">
{{ submitting ? "Создание..." : "Создать канал" }}
</button>
</div>
</form>
<p v-if="error" class="form-error">{{ error }}</p>
</section>
<section class="card">
<h3>Список <span v-if="loading" class="muted">(загрузка...)</span></h3>
<div class="table-wrap">
<table class="admin-table">
<thead>
<tr>
<th>ID</th>
<th>Название</th>
<th>Канал</th>
<th>Тип</th>
<th>Активен</th>
<th>Создан</th>
<th></th>
</tr>
</thead>
<tbody>
<tr v-for="ch in channels" :key="ch.id">
<td>{{ ch.id }}</td>
<td>{{ ch.name }}</td>
<td><code>@{{ ch.channel }}</code></td>
<td>{{ ch.source_type }}</td>
<td>{{ ch.is_active ? "да" : "нет" }}</td>
<td>{{ formatDate(ch.created_at) }}</td>
<td class="actions">
<button type="button" class="btn btn-sm" @click="openEdit(ch)">Изменить</button>
<button type="button" class="btn btn-sm btn-danger" @click="handleDelete(ch)">
Удалить
</button>
</td>
</tr>
<tr v-if="!loading && !channels.length">
<td colspan="7" class="muted">Пока нет каналов</td>
</tr>
</tbody>
</table>
</div>
</section>
<div v-if="editModal.open" class="modal-backdrop" @click.self="closeEdit">
<div class="modal">
<h3>Редактировать канал #{{ editModal.channel?.id }}</h3>
<form class="admin-form" @submit.prevent="handleSaveEdit">
<label>
Название
<input v-model="editModal.name" type="text" required />
</label>
<label>
Handle / id
<input v-model="editModal.handle" type="text" required />
</label>
<label class="checkbox-label">
<input v-model="editModal.is_active" type="checkbox" />
Активен
</label>
<div class="form-actions">
<button type="button" class="btn" @click="closeEdit">Отмена</button>
<button type="submit" class="btn btn-primary" :disabled="submitting">Сохранить</button>
</div>
</form>
</div>
</div>
</div>
</template>
<style scoped>
.intro {
margin: 0 0 1rem;
font-size: 0.9rem;
color: #555;
line-height: 1.5;
}
.checkbox-label {
display: flex;
align-items: center;
gap: 0.5rem;
margin-top: 1.5rem;
}
.actions {
display: flex;
gap: 0.35rem;
flex-wrap: wrap;
}
.modal-backdrop {
position: fixed;
inset: 0;
background: rgba(0, 0, 0, 0.35);
display: flex;
align-items: center;
justify-content: center;
z-index: 50;
}
.modal {
background: #fff;
border-radius: 8px;
padding: 1.25rem 1.5rem;
width: min(420px, 92vw);
box-shadow: 0 8px 24px rgba(0, 0, 0, 0.15);
}
.modal h3 {
margin: 0 0 1rem;
font-size: 1.05rem;
}
.modal .admin-form label {
display: block;
margin-bottom: 0.75rem;
}
</style>
@@ -1,7 +1,7 @@
<script setup lang="ts">
import { computed, onMounted, ref, watch } from "vue";
import { useRouter } from "vue-router";
import { fetchEvents } from "../api/admin";
import { createEvent, deleteEvent, fetchEvents, updateEvent } from "../api/admin";
import type { EventFilters, EventRecord } from "../types/admin";
const router = useRouter();
@@ -10,6 +10,7 @@ const items = ref<EventRecord[]>([]);
const total = ref(0);
const loading = ref(true);
const error = ref("");
const submitting = ref(false);
const filters = ref({
search: "",
@@ -24,8 +25,44 @@ const filters = ref({
const page = ref(1);
const pageSize = 25;
const selectedEvent = ref<EventRecord | null>(null);
const modalOpen = ref(false);
type EventForm = {
source_type: string;
source_url: string;
title: string;
description: string;
locality: string;
region: string;
topic: string;
raw_text: string;
latitude: string;
longitude: string;
event_date: string;
tags: string;
};
function emptyForm(): EventForm {
return {
source_type: "manual",
source_url: "",
title: "",
description: "",
locality: "",
region: "",
topic: "",
raw_text: "",
latitude: "",
longitude: "",
event_date: "",
tags: "",
};
}
const createForm = ref(emptyForm());
const editModal = ref<{ open: boolean; event: EventRecord | null; form: EventForm }>({
open: false,
event: null,
form: emptyForm(),
});
const totalPages = computed(() => Math.max(1, Math.ceil(total.value / pageSize)));
@@ -76,14 +113,128 @@ function resetFilters() {
loadEvents();
}
function openModal(event: EventRecord) {
selectedEvent.value = event;
modalOpen.value = true;
function parseOptionalFloat(value: string): number | null {
const trimmed = value.trim();
if (!trimmed) return null;
const num = Number(trimmed);
if (Number.isNaN(num)) throw new Error("Координаты должны быть числами");
return num;
}
function closeModal() {
modalOpen.value = false;
selectedEvent.value = null;
function parseTags(value: string): string[] | null {
const tags = value
.split(",")
.map((t) => t.trim())
.filter(Boolean);
return tags.length ? tags : null;
}
function toDateTimeLocal(value: string | null): string {
if (!value) return "";
const d = new Date(value);
if (Number.isNaN(d.getTime())) return "";
const pad = (n: number) => String(n).padStart(2, "0");
return `${d.getFullYear()}-${pad(d.getMonth() + 1)}-${pad(d.getDate())}T${pad(d.getHours())}:${pad(d.getMinutes())}`;
}
function formToPayload(form: EventForm) {
const latitude = parseOptionalFloat(form.latitude);
const longitude = parseOptionalFloat(form.longitude);
if ((latitude == null) !== (longitude == null)) {
throw new Error("Укажите обе координаты или ни одной");
}
const source_url = form.source_url.trim();
if (!source_url) throw new Error("Укажите source_url");
return {
source_type: form.source_type.trim() || "manual",
source_url,
title: form.title.trim(),
description: form.description.trim(),
locality: form.locality.trim(),
region: form.region.trim() || null,
topic: form.topic.trim() || null,
raw_text: form.raw_text,
latitude,
longitude,
event_date: form.event_date ? new Date(form.event_date).toISOString() : null,
tags: parseTags(form.tags),
};
}
function eventToForm(event: EventRecord): EventForm {
return {
source_type: event.source_type || "manual",
source_url: event.source_url,
title: event.title || "",
description: event.description || "",
locality: event.locality || "",
region: event.region || "",
topic: event.topic || "",
raw_text: event.raw_text || "",
latitude: event.latitude != null ? String(event.latitude) : "",
longitude: event.longitude != null ? String(event.longitude) : "",
event_date: toDateTimeLocal(event.event_date),
tags: event.tags?.join(", ") ?? "",
};
}
async function handleCreate() {
submitting.value = true;
error.value = "";
try {
await createEvent(formToPayload(createForm.value));
createForm.value = emptyForm();
createForm.value.source_url = `manual://${crypto.randomUUID()}`;
await loadEvents();
} catch (err) {
error.value = err instanceof Error ? err.message : "Не удалось создать событие";
} finally {
submitting.value = false;
}
}
function openEdit(event: EventRecord) {
editModal.value = {
open: true,
event,
form: eventToForm(event),
};
}
function closeEdit() {
editModal.value.open = false;
editModal.value.event = null;
}
async function handleSaveEdit() {
if (!editModal.value.event) return;
submitting.value = true;
error.value = "";
try {
await updateEvent(editModal.value.event.id, formToPayload(editModal.value.form));
closeEdit();
await loadEvents();
} catch (err) {
error.value = err instanceof Error ? err.message : "Не удалось сохранить событие";
} finally {
submitting.value = false;
}
}
async function handleDelete(event: EventRecord) {
const label = event.title || event.source_url || `#${event.id}`;
if (!window.confirm(`Удалить событие #${event.id} («${label}»)? Связанная точка на карте тоже будет удалена.`)) {
return;
}
error.value = "";
try {
await deleteEvent(event.id);
if (editModal.value.event?.id === event.id) closeEdit();
await loadEvents();
} catch (err) {
error.value = err instanceof Error ? err.message : "Не удалось удалить событие";
}
}
function goToMap(event: EventRecord) {
@@ -103,13 +254,77 @@ function truncate(text: string, max = 80): string {
watch(page, loadEvents);
onMounted(loadEvents);
onMounted(() => {
createForm.value.source_url = `manual://${crypto.randomUUID()}`;
loadEvents();
});
</script>
<template>
<div class="page">
<h2 class="page-heading">События</h2>
<section class="card">
<h3>Новое событие</h3>
<form class="admin-form" @submit.prevent="handleCreate">
<div class="form-row">
<label>
Заголовок
<input v-model="createForm.title" type="text" placeholder="Краткий заголовок" />
</label>
<label>
Источник (type)
<input v-model="createForm.source_type" type="text" placeholder="manual" />
</label>
<label class="span-2">
source_url
<input v-model="createForm.source_url" type="text" required />
</label>
<label class="span-2">
Описание
<textarea v-model="createForm.description" rows="2" />
</label>
<label>
Регион
<input v-model="createForm.region" type="text" />
</label>
<label>
Тема
<input v-model="createForm.topic" type="text" />
</label>
<label>
Населённый пункт
<input v-model="createForm.locality" type="text" />
</label>
<label>
Дата события
<input v-model="createForm.event_date" type="datetime-local" />
</label>
<label>
Широта
<input v-model="createForm.latitude" type="text" placeholder="48.1234" />
</label>
<label>
Долгота
<input v-model="createForm.longitude" type="text" placeholder="37.6543" />
</label>
<label>
Теги (через запятую)
<input v-model="createForm.tags" type="text" placeholder="обстрел, нп" />
</label>
<label class="span-2">
Сырой текст
<textarea v-model="createForm.raw_text" rows="3" />
</label>
</div>
<div class="form-actions">
<button type="submit" class="btn btn-primary" :disabled="submitting">
{{ submitting ? "Создание..." : "Создать событие" }}
</button>
</div>
</form>
</section>
<section class="card">
<form class="admin-form filters-form" @submit.prevent="applyFilters">
<div class="form-row">
@@ -177,7 +392,7 @@ onMounted(loadEvents);
<tr v-for="event in items" :key="event.id">
<td>{{ event.id }}</td>
<td>
<button class="link-btn" @click="openModal(event)">
<button class="link-btn" type="button" @click="openEdit(event)">
{{ truncate(event.title || event.description || "—") }}
</button>
</td>
@@ -191,14 +406,19 @@ onMounted(loadEvents);
</template>
<template v-else>—</template>
</td>
<td>
<td class="actions">
<button class="btn btn-sm" type="button" @click="openEdit(event)">Изменить</button>
<button
v-if="event.latitude != null && event.longitude != null"
class="btn btn-sm"
type="button"
@click="goToMap(event)"
>
На карте
</button>
<button class="btn btn-sm btn-danger" type="button" @click="handleDelete(event)">
Удалить
</button>
</td>
</tr>
</tbody>
@@ -206,56 +426,129 @@ onMounted(loadEvents);
</div>
<div class="pagination">
<button class="btn btn-sm" :disabled="page <= 1" @click="page--">← Назад</button>
<button class="btn btn-sm" type="button" :disabled="page <= 1" @click="page--">← Назад</button>
<span>Стр. {{ page }} / {{ totalPages }}</span>
<button class="btn btn-sm" :disabled="page >= totalPages" @click="page++">Вперёд →</button>
<button class="btn btn-sm" type="button" :disabled="page >= totalPages" @click="page++">
Вперёд →
</button>
</div>
</section>
<div v-if="modalOpen && selectedEvent" class="modal-overlay" @click.self="closeModal">
<div class="modal">
<div v-if="editModal.open && editModal.event" class="modal-overlay" @click.self="closeEdit">
<div class="modal modal-wide">
<header class="modal-header">
<h3>Событие #{{ selectedEvent.id }}</h3>
<button class="btn btn-sm" @click="closeModal">✕</button>
<h3>Событие #{{ editModal.event.id }}</h3>
<button class="btn btn-sm" type="button" @click="closeEdit">✕</button>
</header>
<div class="modal-body">
<dl class="detail-list">
<dt>Заголовок</dt>
<dd>{{ selectedEvent.title || "—" }}</dd>
<dt>Описание</dt>
<dd>{{ selectedEvent.description || "—" }}</dd>
<dt>Источник</dt>
<dd>{{ selectedEvent.source_type }} — {{ selectedEvent.source_url }}</dd>
<dt>Регион / Тема</dt>
<dd>{{ selectedEvent.region ?? "—" }} / {{ selectedEvent.topic ?? "—" }}</dd>
<dt>Населённый пункт</dt>
<dd>{{ selectedEvent.locality || "—" }}</dd>
<dt>Дата события</dt>
<dd>{{ formatDate(selectedEvent.event_date) }}</dd>
<dt>Ingested</dt>
<dd>{{ formatDate(selectedEvent.ingested_at) }}</dd>
<dt>Координаты</dt>
<dd>
<template v-if="selectedEvent.latitude != null">
{{ selectedEvent.latitude }}, {{ selectedEvent.longitude }}
</template>
<template v-else>—</template>
</dd>
<dt>Текст</dt>
<dd class="raw-text">{{ selectedEvent.raw_text || "—" }}</dd>
</dl>
</div>
<footer class="modal-footer">
<button
v-if="selectedEvent.latitude != null"
class="btn btn-primary"
@click="goToMap(selectedEvent)"
>
Показать на карте
</button>
<button class="btn" @click="closeModal">Закрыть</button>
</footer>
<form class="modal-body admin-form" @submit.prevent="handleSaveEdit">
<div class="form-row">
<label>
Заголовок
<input v-model="editModal.form.title" type="text" />
</label>
<label>
Источник (type)
<input v-model="editModal.form.source_type" type="text" />
</label>
<label class="span-2">
source_url
<input v-model="editModal.form.source_url" type="text" required />
</label>
<label class="span-2">
Описание
<textarea v-model="editModal.form.description" rows="2" />
</label>
<label>
Регион
<input v-model="editModal.form.region" type="text" />
</label>
<label>
Тема
<input v-model="editModal.form.topic" type="text" />
</label>
<label>
Населённый пункт
<input v-model="editModal.form.locality" type="text" />
</label>
<label>
Дата события
<input v-model="editModal.form.event_date" type="datetime-local" />
</label>
<label>
Широта
<input v-model="editModal.form.latitude" type="text" />
</label>
<label>
Долгота
<input v-model="editModal.form.longitude" type="text" />
</label>
<label>
Теги
<input v-model="editModal.form.tags" type="text" />
</label>
<label class="span-2">
Сырой текст
<textarea v-model="editModal.form.raw_text" rows="4" />
</label>
<p class="hint span-2">
Ingested: {{ formatDate(editModal.event.ingested_at) }}. При координатах создаётся/обновляется
точка на карте; без координат связанная точка удаляется.
</p>
</div>
<div class="form-actions">
<button type="button" class="btn" @click="closeEdit">Отмена</button>
<button
v-if="editModal.event.latitude != null"
type="button"
class="btn"
@click="goToMap(editModal.event)"
>
На карте
</button>
<button type="submit" class="btn btn-primary" :disabled="submitting">
{{ submitting ? "Сохранение..." : "Сохранить" }}
</button>
</div>
</form>
</div>
</div>
</div>
</template>
<style scoped>
.form-row .span-2 {
grid-column: span 2;
}
.form-row textarea {
width: 100%;
font: inherit;
padding: 0.4rem 0.5rem;
border: 1px solid #d1d5db;
border-radius: 4px;
resize: vertical;
}
.actions {
display: flex;
gap: 0.35rem;
flex-wrap: wrap;
white-space: nowrap;
}
.hint {
margin: 0;
font-size: 0.85rem;
color: #666;
line-height: 1.4;
}
.modal-wide {
width: min(720px, 94vw);
}
.mono {
font-family: ui-monospace, SFMono-Regular, Menlo, Consolas, monospace;
font-size: 0.8rem;
}
</style>
@@ -0,0 +1,146 @@
<script setup lang="ts">
import { ref } from "vue";
import { useRoute, useRouter } from "vue-router";
import { useAuth } from "../composables/useAuth";
const route = useRoute();
const router = useRouter();
const { login, busy, error, authenticated } = useAuth();
const username = ref("admin");
const password = ref("");
if (authenticated.value) {
const next = typeof route.query.next === "string" ? route.query.next : "/";
void router.replace(next || "/");
}
async function onSubmit() {
const ok = await login(username.value, password.value);
if (!ok) return;
const next = typeof route.query.next === "string" ? route.query.next : "/";
await router.replace(next || "/");
}
</script>
<template>
<div class="login-page">
<form class="login-card" @submit.prevent="onSubmit">
<h1>Вход в админку</h1>
<p class="hint">Карта доступна без входа. Остальные разделы — после авторизации.</p>
<label>
Логин
<input v-model="username" type="text" autocomplete="username" required />
</label>
<label>
Пароль
<input v-model="password" type="password" autocomplete="current-password" required />
</label>
<p v-if="error" class="error">{{ error }}</p>
<button type="submit" class="btn-primary" :disabled="busy">
{{ busy ? "Вход..." : "Войти" }}
</button>
<router-link class="back-link" to="/">← К карте</router-link>
</form>
</div>
</template>
<style scoped>
.login-page {
display: flex;
align-items: center;
justify-content: center;
min-height: 100%;
padding: 2rem 1rem;
background: linear-gradient(160deg, #eef2f7 0%, #f5f5f5 45%, #e8edf3 100%);
}
.login-card {
width: 100%;
max-width: 380px;
display: flex;
flex-direction: column;
gap: 0.85rem;
padding: 1.75rem;
background: #fff;
border: 1px solid #e0e0e0;
border-radius: 10px;
box-shadow: 0 8px 24px rgba(0, 0, 0, 0.06);
}
h1 {
margin: 0;
font-size: 1.25rem;
font-weight: 650;
}
.hint {
margin: 0;
font-size: 0.875rem;
color: #666;
line-height: 1.4;
}
label {
display: flex;
flex-direction: column;
gap: 0.35rem;
font-size: 0.875rem;
color: #444;
}
input {
padding: 0.55rem 0.7rem;
border: 1px solid #d0d0d0;
border-radius: 6px;
background: #fff;
}
input:focus {
outline: 2px solid #bbdefb;
border-color: #64b5f6;
}
.error {
margin: 0;
font-size: 0.85rem;
color: #b91c1c;
}
.btn-primary {
margin-top: 0.25rem;
padding: 0.65rem 1rem;
border: none;
border-radius: 6px;
background: #1565c0;
color: #fff;
font-weight: 500;
cursor: pointer;
}
.btn-primary:disabled {
opacity: 0.65;
cursor: not-allowed;
}
.btn-primary:not(:disabled):hover {
background: #0d47a1;
}
.back-link {
text-align: center;
font-size: 0.875rem;
color: #1565c0;
text-decoration: none;
}
.back-link:hover {
text-decoration: underline;
}
</style>
@@ -16,12 +16,14 @@ import CoordsTools from "../components/map/CoordsTools.vue";
import MapToolbar from "../components/map/MapToolbar.vue";
import PlaceSearch from "../components/map/PlaceSearch.vue";
import ObjectPanel from "../components/ObjectPanel.vue";
import { useAuth } from "../composables/useAuth";
import type { LeafletMapApi } from "../composables/useLeafletMap";
import type { DateRangePreset, MapFilters, MapObjectWithEvent, MapQueryParams } from "../types/map";
import type { MapObjectCreate, ObjectType } from "../types/object";
const route = useRoute();
const router = useRouter();
const { authenticated } = useAuth();
const objects = ref<MapObjectWithEvent[]>([]);
const mapFilters = ref<MapFilters | null>(null);
@@ -162,6 +164,7 @@ function handleMapContextMenu(payload: {
longitude: number;
object: MapObjectWithEvent | null;
}) {
if (!authenticated.value) return;
contextMenu.value = {
visible: true,
x: payload.x,
@@ -257,6 +260,7 @@ async function handleMoveObject(payload: {
latitude: number;
longitude: number;
}) {
if (!authenticated.value) return;
if (payload.object.event_id) return;
try {
await updateObject(payload.object.id, {
@@ -327,17 +331,22 @@ onUnmounted(() => {
:objects="objects"
:selected-id="selectedId"
:open-popup-id="openPopupId"
:editable="authenticated"
@ready="onMapReady"
@center-change="onCenterChange"
@select="handleSelectObject"
@contextmenu="handleMapContextMenu"
@move="handleMoveObject"
/>
<ObjectPanel :object="selectedObject" @open-events="openInEvents" />
<ObjectPanel
:object="selectedObject"
:can-edit="authenticated"
@open-events="openInEvents"
/>
</div>
<ContextMenu
v-if="contextMenu.visible"
v-if="contextMenu.visible && authenticated"
:x="contextMenu.x"
:y="contextMenu.y"
:target="contextMenu.target"
@@ -350,7 +359,7 @@ onUnmounted(() => {
/>
<CreateObjectModal
v-if="createModal.visible"
v-if="createModal.visible && authenticated"
:latitude="createModal.latitude"
:longitude="createModal.longitude"
@close="closeCreateModal"
@@ -358,7 +367,7 @@ onUnmounted(() => {
/>
<EditObjectModal
v-if="editModal.visible && editModal.object"
v-if="editModal.visible && editModal.object && authenticated"
:object="editModal.object"
@close="closeEditModal"
@submit="handleEditObject"
@@ -0,0 +1,887 @@
<script setup lang="ts">
import { onMounted, ref, watch } from "vue";
import {
createProfile,
deleteProfile,
fetchLlmDefaults,
fetchProfiles,
fetchTargetFields,
generateParserProfile,
previewLlmProfile,
previewParserProfile,
updateProfile,
type HeuristicProfile,
type LlmProfile,
type TargetField,
} from "../api/entities";
import type { ParserProfile } from "../types/admin";
const profiles = ref<ParserProfile[]>([]);
const loading = ref(true);
const error = ref("");
const success = ref("");
const submitting = ref(false);
const profileKind = ref<"heuristic" | "llm">("heuristic");
const samplePost = ref("");
const generateHint = ref("");
const profileName = ref("");
const targetFields = ref<TargetField[]>([]);
const profileJson = ref("");
const llmInstruction = ref("");
const llmSchemaJson = ref("");
const llmMultiEvent = ref(false);
const previewFields = ref<Record<string, string> | null>(null);
const previewEvents = ref<Record<string, string>[]>([]);
const emptyFields = ref<string[]>([]);
const requiredFields = ref<string[]>([]);
const previewMatched = ref(true);
const previewMissing = ref<string[]>([]);
const previewIsEvent = ref(true);
const previewMatchedCount = ref(0);
const generating = ref(false);
const previewing = ref(false);
const loadingFields = ref(true);
const editingId = ref<number | null>(null);
const profileStatus = ref<"draft" | "ready">("draft");
const llmDefaults = ref<LlmProfile | null>(null);
async function loadProfiles() {
loading.value = true;
try {
profiles.value = await fetchProfiles();
error.value = "";
} catch (err) {
error.value = err instanceof Error ? err.message : "Не удалось загрузить профили";
} finally {
loading.value = false;
}
}
async function loadTargetFields() {
loadingFields.value = true;
try {
const data = await fetchTargetFields();
targetFields.value = data.fields;
} catch (err) {
error.value = err instanceof Error ? err.message : "Не удалось загрузить целевые поля";
} finally {
loadingFields.value = false;
}
}
async function loadLlmDefaults() {
try {
llmDefaults.value = await fetchLlmDefaults();
} catch {
llmDefaults.value = null;
}
}
function applyLlmDefaults() {
const d = llmDefaults.value;
if (!d) return;
llmInstruction.value = d.instruction || "";
llmSchemaJson.value = JSON.stringify(d.extract_schema || {}, null, 2);
if (!requiredFields.value.length) {
requiredFields.value = Array.isArray(d.required_fields) ? [...d.required_fields] : [];
}
if (typeof d.multi_event === "boolean") {
llmMultiEvent.value = d.multi_event;
}
}
function syncHeuristicFromJson(): HeuristicProfile | null {
if (!profileJson.value.trim()) return null;
try {
return JSON.parse(profileJson.value) as HeuristicProfile;
} catch {
error.value = "Некорректный JSON профиля";
return null;
}
}
function syncLlmFromForm(): LlmProfile | null {
let schema: Record<string, string>;
try {
schema = JSON.parse(llmSchemaJson.value || "{}") as Record<string, string>;
} catch {
error.value = "Некорректный JSON extract_schema";
return null;
}
if (!schema || typeof schema !== "object" || !Object.keys(schema).length) {
error.value = "extract_schema не должен быть пустым";
return null;
}
return {
instruction: llmInstruction.value.trim() || null,
extract_schema: schema,
required_fields: [...requiredFields.value],
multi_event: llmMultiEvent.value,
};
}
const DEFAULT_REQUIRED = ["event_date", "coords"];
function uniqueFields(names: string[]): string[] {
return names.filter((name, idx) => names.indexOf(name) === idx);
}
function writeRequiredToHeuristicJson(names: string[]) {
const current = syncHeuristicFromJson();
if (!current) return;
current.required_fields = names;
profileJson.value = JSON.stringify(current, null, 2);
}
function requiredFromHeuristic(
profile: HeuristicProfile,
preview: Record<string, string> | null,
): string[] {
if (Array.isArray(profile.required_fields)) {
return uniqueFields(profile.required_fields);
}
if (!preview) return [];
return DEFAULT_REQUIRED.filter((name) => !!(preview[name] || "").trim());
}
function isRequired(name: string): boolean {
return requiredFields.value.includes(name);
}
async function toggleRequired(name: string) {
const next = isRequired(name)
? requiredFields.value.filter((n) => n !== name)
: [...requiredFields.value, name];
requiredFields.value = next;
if (profileKind.value === "heuristic") {
writeRequiredToHeuristicJson(next);
}
if (previewFields.value && samplePost.value.trim()) {
await handlePreview();
}
}
watch(profileKind, (kind, prev) => {
if (kind === prev) return;
previewFields.value = null;
previewEvents.value = [];
emptyFields.value = [];
previewMatched.value = true;
previewMissing.value = [];
previewIsEvent.value = true;
previewMatchedCount.value = 0;
success.value = "";
if (kind === "llm" && !llmSchemaJson.value.trim()) {
applyLlmDefaults();
}
});
async function handleGenerate() {
if (profileKind.value === "llm") {
applyLlmDefaults();
success.value = "Подставлены дефолты LLM schema/instruction. Отредактируйте и Preview.";
return;
}
if (!samplePost.value.trim()) {
error.value = "Вставьте образец поста";
return;
}
generating.value = true;
error.value = "";
success.value = "";
previewFields.value = null;
emptyFields.value = [];
try {
let current: HeuristicProfile | null = null;
if (profileJson.value.trim()) {
current = syncHeuristicFromJson();
if (!current) return;
}
const data = await generateParserProfile({
sample_post: samplePost.value,
hint: generateHint.value.trim() || undefined,
current_profile: current,
});
const required =
current && Array.isArray(current.required_fields)
? uniqueFields(current.required_fields)
: DEFAULT_REQUIRED.filter((name) => !!(data.preview[name] || "").trim());
data.profile.required_fields = required;
profileJson.value = JSON.stringify(data.profile, null, 2);
requiredFields.value = required;
previewFields.value = data.preview;
emptyFields.value = data.empty_fields || [];
const previewData = await previewParserProfile(samplePost.value, data.profile);
previewMatched.value = previewData.matched;
previewMissing.value = previewData.missing_required || [];
if (emptyFields.value.length) {
success.value =
`Профиль обновлён, но пустые поля: ${emptyFields.value.join(", ")}. ` +
"Допишите подсказку ниже и снова нажмите Generate.";
if (!generateHint.value.trim() && emptyFields.value.includes("description")) {
generateHint.value =
"description должен содержать многострочный текст тела поста после заголовка и до координат (47.…, 36.…), без строки заголовка и без хвоста Источник/Геопривязка/#теги";
}
} else {
success.value = "Профиль сгенерирован. Preview и runtime — только по правилам, без LLM.";
}
} catch (err) {
error.value = err instanceof Error ? err.message : "Ошибка генерации";
} finally {
generating.value = false;
}
}
async function handlePreview() {
if (!samplePost.value.trim()) {
error.value = "Нужен образец поста для preview";
return;
}
previewing.value = true;
error.value = "";
try {
if (profileKind.value === "llm") {
const llm = syncLlmFromForm();
if (!llm) return;
const data = await previewLlmProfile(samplePost.value, llm);
const rows =
Array.isArray(data.events) && data.events.length
? data.events
: data.fields
? [data.fields]
: [];
previewEvents.value = rows;
previewFields.value = rows[0] || data.fields || null;
emptyFields.value = previewFields.value
? Object.entries(previewFields.value)
.filter(([, v]) => !v)
.map(([k]) => k)
: [];
previewMatched.value = data.matched;
previewMissing.value = data.missing_required || [];
previewIsEvent.value = data.is_event;
previewMatchedCount.value = data.matched_count ?? rows.length;
if (!data.is_event) {
success.value = "Модель пометила текст как не-событие (is_event=false) — runtime пропустит.";
} else if (rows.length > 1) {
success.value = `Preview: ${rows.length} событий из поста (пройдут фильтр: ${previewMatchedCount.value}).`;
}
} else {
const current = syncHeuristicFromJson();
if (!current) {
if (!error.value) error.value = "Сначала сгенерируйте или вставьте профиль";
return;
}
const data = await previewParserProfile(samplePost.value, current);
previewFields.value = data.fields;
emptyFields.value = Object.entries(data.fields)
.filter(([, v]) => !v)
.map(([k]) => k);
requiredFields.value = requiredFromHeuristic(current, data.fields);
if (!Array.isArray(current.required_fields)) {
writeRequiredToHeuristicJson(requiredFields.value);
}
previewMatched.value = data.matched;
previewMissing.value = data.missing_required || [];
previewIsEvent.value = true;
previewEvents.value = data.fields ? [data.fields] : [];
previewMatchedCount.value = data.matched ? 1 : 0;
}
} catch (err) {
error.value = err instanceof Error ? err.message : "Ошибка preview";
} finally {
previewing.value = false;
}
}
function resetEditor() {
editingId.value = null;
profileKind.value = "heuristic";
profileName.value = "";
samplePost.value = "";
generateHint.value = "";
profileJson.value = "";
llmInstruction.value = "";
llmSchemaJson.value = "";
llmMultiEvent.value = false;
profileStatus.value = "draft";
previewFields.value = null;
previewEvents.value = [];
emptyFields.value = [];
requiredFields.value = [];
previewMatched.value = true;
previewMissing.value = [];
previewIsEvent.value = true;
previewMatchedCount.value = 0;
success.value = "";
}
function openEdit(p: ParserProfile) {
editingId.value = p.id;
profileKind.value = p.kind === "llm" ? "llm" : "heuristic";
profileName.value = p.name;
samplePost.value = p.sample_post || "";
generateHint.value = "";
profileStatus.value = p.status === "ready" ? "ready" : "draft";
previewFields.value = null;
previewEvents.value = [];
emptyFields.value = [];
previewMatched.value = true;
previewMissing.value = [];
previewIsEvent.value = true;
previewMatchedCount.value = 0;
success.value = "";
error.value = "";
if (p.kind === "llm") {
profileJson.value = "";
const lp = (p.llm_profile || {}) as LlmProfile;
llmInstruction.value = lp.instruction || "";
llmSchemaJson.value = JSON.stringify(lp.extract_schema || {}, null, 2);
llmMultiEvent.value = !!lp.multi_event;
requiredFields.value = Array.isArray(lp.required_fields)
? uniqueFields(lp.required_fields)
: [];
if (!llmSchemaJson.value.trim() || llmSchemaJson.value === "{}") {
applyLlmDefaults();
}
} else {
llmInstruction.value = "";
llmSchemaJson.value = "";
llmMultiEvent.value = false;
profileJson.value = p.heuristic_profile
? JSON.stringify(p.heuristic_profile, null, 2)
: "";
const hp = p.heuristic_profile;
if (hp && Array.isArray(hp.required_fields)) {
requiredFields.value = uniqueFields(hp.required_fields as string[]);
} else if (hp) {
requiredFields.value = [...DEFAULT_REQUIRED];
writeRequiredToHeuristicJson(requiredFields.value);
} else {
requiredFields.value = [];
}
}
window.scrollTo({ top: 0, behavior: "smooth" });
}
async function handleSave() {
if (!profileName.value.trim()) {
error.value = "Укажите название профиля";
return;
}
submitting.value = true;
error.value = "";
success.value = "";
try {
if (profileKind.value === "llm") {
const llm = syncLlmFromForm();
if (!llm) return;
const payload = {
name: profileName.value.trim(),
kind: "llm" as const,
sample_post: samplePost.value,
llm_profile: llm,
heuristic_profile: null,
status: profileStatus.value,
};
if (editingId.value != null) {
await updateProfile(editingId.value, payload);
success.value = `LLM-профиль #${editingId.value} обновлён`;
} else {
const created = await createProfile({
...payload,
status: profileStatus.value === "draft" ? "draft" : "ready",
});
success.value = `LLM-профиль #${created.id} сохранён`;
editingId.value = created.id;
profileStatus.value = created.status === "ready" ? "ready" : "draft";
}
} else {
writeRequiredToHeuristicJson(requiredFields.value);
const current = syncHeuristicFromJson();
if (!current) {
if (!error.value) error.value = "Нужен JSON профиля (Generate или вручную)";
return;
}
const payload = {
name: profileName.value.trim(),
kind: "heuristic" as const,
sample_post: samplePost.value,
heuristic_profile: current,
llm_profile: null,
status: profileStatus.value,
};
if (editingId.value != null) {
await updateProfile(editingId.value, payload);
success.value = `Профиль #${editingId.value} обновлён`;
} else {
const created = await createProfile({
...payload,
status: profileStatus.value === "draft" ? "draft" : "ready",
});
success.value = `Профиль #${created.id} сохранён`;
editingId.value = created.id;
profileStatus.value = created.status === "ready" ? "ready" : "draft";
}
}
await loadProfiles();
} catch (err) {
error.value = err instanceof Error ? err.message : "Не удалось сохранить профиль";
} finally {
submitting.value = false;
}
}
async function handleDelete(p: ParserProfile) {
if (!window.confirm(`Удалить профиль «${p.name}»?`)) return;
error.value = "";
try {
await deleteProfile(p.id);
if (editingId.value === p.id) resetEditor();
await loadProfiles();
} catch (err) {
error.value = err instanceof Error ? err.message : "Не удалось удалить";
}
}
function formatDate(value: string): string {
return new Date(value).toLocaleString("ru-RU");
}
function kindLabel(kind: string | undefined): string {
return kind === "llm" ? "llm" : "heuristic";
}
function canSave(): boolean {
if (!profileName.value.trim()) return false;
if (profileKind.value === "llm") return !!llmSchemaJson.value.trim();
return !!profileJson.value.trim();
}
onMounted(async () => {
await Promise.all([loadProfiles(), loadTargetFields(), loadLlmDefaults()]);
});
</script>
<template>
<div class="page">
<h2 class="page-heading">Профили парсера</h2>
<p class="intro">
Два вида профилей под фиксированные поля Event:
<strong>heuristic</strong> — статичные правила (Generate один раз, runtime без LLM);
<strong>llm</strong> — instruction + schema, DeepSeek на каждый пост в batch.
Связка с каналом — на
<router-link to="/parsers">Парсеры</router-link>.
</p>
<section class="card">
<div class="card-head">
<h3>{{ editingId != null ? `Редактирование #${editingId}` : "Новый профиль" }}</h3>
<button v-if="editingId != null" type="button" class="btn btn-sm" @click="resetEditor">
Сбросить
</button>
</div>
<div class="form-row">
<label>
Название
<input v-model="profileName" type="text" placeholder="Сводка: дата + НП + coords" />
</label>
<label>
Тип
<select v-model="profileKind" :disabled="editingId != null">
<option value="heuristic">heuristic (правила)</option>
<option value="llm">llm (DeepSeek runtime)</option>
</select>
</label>
<label>
Статус
<select v-model="profileStatus">
<option value="draft">draft</option>
<option value="ready">ready</option>
</select>
</label>
</div>
<label>
Образец поста
<textarea
v-model="samplePost"
class="sample-area"
rows="10"
placeholder="15.03.2024 Населённый пункт&#10;&#10;Описание…&#10;&#10;48.123456, 37.654321"
/>
</label>
<template v-if="profileKind === 'heuristic'">
<label>
Подсказка агенту
<span class="muted"> (если Generate ошибся — опишите, что исправить, и нажмите Generate снова)</span>
<textarea
v-model="generateHint"
class="hint-area"
rows="3"
placeholder="например: description — абзацы между заголовком и координатами, без #хештегов"
/>
</label>
</template>
<template v-else>
<label>
Instruction
<textarea
v-model="llmInstruction"
class="hint-area"
rows="3"
placeholder="Инструкция для DeepSeek на каждый пост"
/>
</label>
<label class="required-item multi-flag">
<input v-model="llmMultiEvent" type="checkbox" />
Несколько событий в посте
<span class="muted"> (LLM вернёт массив; ingest с URL #e1, #e2…)</span>
</label>
</template>
<div class="form-actions">
<button
type="button"
class="btn btn-primary"
:disabled="generating || (profileKind === 'heuristic' && !samplePost.trim())"
@click="handleGenerate"
>
<template v-if="profileKind === 'llm'">
{{ generating ? "…" : "Подставить дефолты schema" }}
</template>
<template v-else>
{{
generating
? "Генерация..."
: generateHint.trim() || profileJson.trim()
? "Generate / Refine"
: "Generate"
}}
</template>
</button>
<button
type="button"
class="btn"
:disabled="
previewing ||
(profileKind === 'heuristic' ? !profileJson.trim() : !llmSchemaJson.trim())
"
@click="handlePreview"
>
{{ previewing ? "Preview..." : "Preview" }}
</button>
<button
type="button"
class="btn btn-primary"
:disabled="submitting || !canSave()"
@click="handleSave"
>
{{ submitting ? "Сохранение..." : editingId != null ? "Обновить" : "Сохранить профиль" }}
</button>
</div>
</section>
<div class="grid-2">
<section class="card">
<h3>Целевые поля <span class="muted">(Event)</span></h3>
<p v-if="loadingFields" class="muted">Загрузка…</p>
<div v-else class="table-wrap">
<table class="admin-table">
<thead>
<tr>
<th>Поле</th>
<th>Тип</th>
<th>Описание</th>
</tr>
</thead>
<tbody>
<tr v-for="field in targetFields" :key="field.name">
<td><code>{{ field.name }}</code></td>
<td>{{ field.type }}</td>
<td>{{ field.description }}</td>
</tr>
</tbody>
</table>
</div>
</section>
<section class="card">
<h3 v-if="profileKind === 'heuristic'">Профиль (JSON)</h3>
<h3 v-else>extract_schema (JSON)</h3>
<textarea
v-if="profileKind === 'heuristic'"
v-model="profileJson"
class="profile-area"
rows="16"
spellcheck="false"
placeholder="Появится после Generate"
/>
<textarea
v-else
v-model="llmSchemaJson"
class="profile-area"
rows="16"
spellcheck="false"
placeholder="Ключ → описание поля для LLM"
/>
</section>
</div>
<section v-if="previewFields" class="card">
<h3>
Preview
<span v-if="previewEvents.length > 1" class="muted">
({{ previewEvents.length }} событий)
</span>
<span v-else>строки</span>
</h3>
<p v-if="profileKind === 'llm' && !previewIsEvent" class="form-warn">
is_event=false — пост не будет ингеститься.
</p>
<p v-if="profileKind === 'llm' && previewEvents.length > 1" class="form-success match-ok">
Из поста извлечено {{ previewEvents.length }} событий;
пройдут required_fields: {{ previewMatchedCount }}.
Runtime URL: пост#e1 … #e{{ previewEvents.length }}.
</p>
<p v-if="emptyFields.length" class="form-warn">
Пустые поля (первая строка): <code>{{ emptyFields.join(", ") }}</code>
<template v-if="profileKind === 'heuristic'">
— уточните подсказку и Generate / Refine.
</template>
</p>
<p v-if="!requiredFields.length" class="form-warn">
Нет обязательных полей —
<template v-if="profileKind === 'llm'">
фильтр только по is_event.
</template>
<template v-else>
любой пост канала будет сохранён. Отметьте поля, без которых пост нужно отбрасывать.
</template>
</p>
<p v-else-if="!previewMatched && previewEvents.length <= 1" class="form-warn">
Этот образец не прошёл бы фильтр. Не хватает:
<code>{{ previewMissing.join(", ") }}</code>
</p>
<p v-else-if="previewMatched" class="form-success match-ok">
<template v-if="previewEvents.length <= 1">
Образец проходит фильтр. Runtime сохранит только посты с заполненными
<code>{{ requiredFields.join(", ") }}</code>.
</template>
<template v-else>
Обязательные поля: <code>{{ requiredFields.join(", ") }}</code>
(проверка на каждое событие).
</template>
</p>
<div class="table-wrap">
<table class="admin-table">
<thead>
<tr>
<th v-if="previewEvents.length > 1">#</th>
<th v-for="key in Object.keys(previewFields)" :key="key">{{ key }}</th>
</tr>
</thead>
<tbody>
<tr v-for="(row, idx) in (previewEvents.length ? previewEvents : [previewFields])" :key="idx">
<td v-if="previewEvents.length > 1">{{ idx + 1 }}</td>
<td
v-for="key in Object.keys(previewFields)"
:key="key"
class="preview-cell"
:class="{ empty: !row?.[key] }"
>
{{ row?.[key] || "—" }}
</td>
</tr>
</tbody>
</table>
</div>
<div class="required-box">
<h4>Обязательно для сохранения</h4>
<p class="muted required-hint">
<template v-if="profileKind === 'llm'">
После LLM: событие отбрасывается, если обязательное поле пустое (поверх is_event).
</template>
<template v-else>
Пост отбрасывается, если хоть одно отмеченное поле не извлеклось (для coords — ещё и не парсится lat,lon).
</template>
</p>
<div class="required-grid">
<label
v-for="key in Object.keys(previewFields)"
:key="key"
class="required-item"
>
<input
type="checkbox"
:checked="isRequired(key)"
@change="toggleRequired(key)"
/>
<code>{{ key }}</code>
</label>
</div>
</div>
</section>
<p v-if="error" class="form-error">{{ error }}</p>
<p v-if="success" class="form-success">{{ success }}</p>
<section class="card">
<h3>Сохранённые профили <span v-if="loading" class="muted">(загрузка...)</span></h3>
<div class="table-wrap">
<table class="admin-table">
<thead>
<tr>
<th>ID</th>
<th>Название</th>
<th>Тип</th>
<th>Статус</th>
<th>Создан</th>
<th></th>
</tr>
</thead>
<tbody>
<tr v-for="p in profiles" :key="p.id">
<td>{{ p.id }}</td>
<td>{{ p.name }}</td>
<td><code>{{ kindLabel(p.kind) }}</code></td>
<td>{{ p.status }}</td>
<td>{{ formatDate(p.created_at) }}</td>
<td class="actions">
<button type="button" class="btn btn-sm" @click="openEdit(p)">Изменить</button>
<button type="button" class="btn btn-sm btn-danger" @click="handleDelete(p)">
Удалить
</button>
</td>
</tr>
<tr v-if="!loading && !profiles.length">
<td colspan="6" class="muted">Пока нет профилей</td>
</tr>
</tbody>
</table>
</div>
</section>
</div>
</template>
<style scoped>
.intro {
margin: 0 0 1rem;
font-size: 0.9rem;
color: #555;
line-height: 1.5;
max-width: 72ch;
}
.card-head {
display: flex;
align-items: center;
justify-content: space-between;
gap: 0.75rem;
margin-bottom: 0.5rem;
}
.card-head h3 {
margin: 0;
}
.sample-area,
.profile-area,
.hint-area {
width: 100%;
font: inherit;
font-family: ui-monospace, SFMono-Regular, Menlo, Consolas, monospace;
font-size: 0.8rem;
padding: 0.6rem 0.75rem;
border: 1px solid #d1d5db;
border-radius: 6px;
resize: vertical;
line-height: 1.45;
margin-top: 0.35rem;
}
.hint-area {
font-family: inherit;
}
.grid-2 {
display: grid;
grid-template-columns: 1fr 1fr;
gap: 1rem;
}
@media (max-width: 960px) {
.grid-2 {
grid-template-columns: 1fr;
}
}
.preview-cell {
max-width: 180px;
white-space: pre-wrap;
word-break: break-word;
font-size: 0.8rem;
}
.preview-cell.empty {
background: #fff8e1;
color: #b26a00;
}
.form-success {
color: #2e7d32;
font-size: 0.875rem;
margin-top: 0.5rem;
}
.form-warn {
color: #b26a00;
font-size: 0.875rem;
margin: 0 0 0.75rem;
}
.match-ok {
margin: 0 0 0.75rem;
}
.required-box {
margin-top: 1rem;
}
.required-box h4 {
margin: 0 0 0.25rem;
font-size: 0.9rem;
}
.required-hint {
margin: 0 0 0.6rem;
font-size: 0.8rem;
}
.required-grid {
display: flex;
flex-wrap: wrap;
gap: 0.5rem 1rem;
}
.required-item {
display: flex;
align-items: center;
gap: 0.35rem;
font-size: 0.85rem;
}
.multi-flag {
margin: 0.75rem 0 0;
flex-wrap: wrap;
}
.actions {
display: flex;
gap: 0.35rem;
flex-wrap: wrap;
}
</style>
@@ -1,7 +1,8 @@
<script setup lang="ts">
import { computed, onMounted, onUnmounted, ref } from "vue";
import { createJob, deleteJob, fetchJobs, retryJob, updateJob } from "../api/admin";
import type { ParseJob } from "../types/admin";
import { fetchChannels, fetchProfiles } from "../api/entities";
import type { ParseChannel, ParseJob, ParserProfile } from "../types/admin";
const INTERVAL_OPTIONS = [
{ value: 900, label: "15 мин" },
@@ -12,83 +13,69 @@ const INTERVAL_OPTIONS = [
{ value: 86400, label: "24 ч" },
];
const SOURCE_TYPES = [
{ value: "telegram", label: "Telegram" },
{ value: "crawl4ai", label: "Crawl4AI (web)" },
{ value: "viina", label: "VIINA (news NLP)" },
] as const;
const jobs = ref<ParseJob[]>([]);
const channels = ref<ParseChannel[]>([]);
const profiles = ref<ParserProfile[]>([]);
const loading = ref(true);
const error = ref("");
const submitting = ref(false);
const forcingJobId = ref<number | null>(null);
const editingJob = ref<ParseJob | null>(null);
const pairForm = ref({
channel_id: null as number | null,
profile_id: null as number | null,
limit: 10,
interval_seconds: 3600,
});
const editForm = ref({
channel: "",
limit: 50,
urlsText: "",
textsText: "",
domain_profile: "generic_news",
input_mode: "urls" as "urls" | "texts" | "mixed",
extract_mode: "heuristic" as "heuristic" | "llm",
channel_id: null as number | null,
profile_id: null as number | null,
limit: 10,
interval_seconds: 3600,
is_active: true,
});
const form = ref({
source_type: "telegram" as "telegram" | "crawl4ai" | "viina",
channel: "",
limit: 50,
urlsText: "",
textsText: "",
domain_profile: "generic_news",
input_mode: "urls" as "urls" | "texts" | "mixed",
extract_mode: "heuristic" as "heuristic" | "llm",
interval_seconds: 3600,
const readyProfiles = computed(() =>
profiles.value.filter((p) => {
if (p.status !== "ready") return false;
if (p.kind === "llm") return !!p.llm_profile;
return !!p.heuristic_profile;
}),
);
const activeChannels = computed(() => channels.value.filter((c) => c.is_active));
const editChannelOptions = computed(() => channels.value);
const editProfileOptions = computed(() => {
const list = [...readyProfiles.value];
const id = editForm.value.profile_id;
if (id != null && !list.some((p) => p.id === id)) {
const found = profiles.value.find((p) => p.id === id);
if (found) list.unshift(found);
}
return list;
});
let pollTimer: ReturnType<typeof setInterval> | null = null;
function channelFromConfig(config: Record<string, unknown>): string {
return typeof config.channel === "string" ? config.channel : "";
function isPairJob(job: ParseJob | null): boolean {
return !!job && job.profile_id != null && job.channel_id != null;
}
function limitFromConfig(config: Record<string, unknown>): number {
const limit = config.limit;
return typeof limit === "number" ? limit : 50;
}
function urlsFromConfig(config: Record<string, unknown>): string[] {
return Array.isArray(config.urls)
? config.urls.filter((u): u is string => typeof u === "string")
: [];
}
function textsFromConfig(config: Record<string, unknown>): string[] {
return Array.isArray(config.texts)
? config.texts.filter((t): t is string => typeof t === "string")
: [];
}
function parseLines(text: string): string[] {
return text
.split(/\r?\n/)
.map((line) => line.trim())
.filter(Boolean);
return typeof limit === "number" ? limit : 10;
}
function sourceSummary(job: ParseJob): string {
const cfg = job.source_config;
const mode = cfg.extract_mode === "llm" ? " [LLM]" : "";
if (job.source_type === "telegram") {
return (channelFromConfig(cfg) || "—") + mode;
if (isPairJob(job)) {
const ch = job.channel_name || job.channel_handle || `#${job.channel_id}`;
const pr = job.profile_name || `#${job.profile_id}`;
return `${ch} × ${pr}`;
}
const urls = urlsFromConfig(cfg);
if (urls.length) return (urls.length === 1 ? urls[0] : `${urls.length} URL`) + mode;
const texts = textsFromConfig(cfg);
if (texts.length) return `${texts.length} текст(ов)`;
return "—";
return `${job.source_type} (без связки)`;
}
function formatInterval(seconds: number): string {
@@ -99,53 +86,6 @@ function formatInterval(seconds: number): string {
return `${Math.round(seconds / 86400)} д`;
}
function buildSourceConfig(
sourceType: string,
data: {
channel: string;
limit: number;
urlsText: string;
textsText: string;
domain_profile: string;
input_mode: "urls" | "texts" | "mixed";
extract_mode: "heuristic" | "llm";
},
): Record<string, unknown> {
if (sourceType === "telegram") {
return {
channel: data.channel.trim(),
limit: data.limit,
extract_mode: data.extract_mode,
};
}
if (sourceType === "crawl4ai") {
return {
urls: parseLines(data.urlsText),
domain_profile: data.domain_profile.trim() || "generic_news",
extract_mode: data.extract_mode,
};
}
return {
urls: parseLines(data.urlsText),
texts: parseLines(data.textsText),
input_mode: data.input_mode,
};
}
const formHint = computed(() => {
if (form.value.source_type === "telegram") {
return form.value.extract_mode === "llm"
? "LLM (DeepSeek): неструктурированные посты канала разбираются в locality/date/coords/description."
: "Активные Telegram-парсеры подхватываются real-time listener; batch забирает последние N постов (эвристики).";
}
if (form.value.source_type === "crawl4ai") {
return form.value.extract_mode === "llm"
? "Crawl4AI + LLM (DeepSeek): страница → structured event. Нужен DEEPSEEK_API_KEY и cp-workers-web."
: "Crawl4AI обходит URL и нормализует страницы эвристиками. Обрабатывает cp-workers-web.";
}
return "VIINA-адаптер извлекает инциденты из новостных URL/текстов. Обрабатывает cp-workers-nlp.";
});
async function loadJobs() {
try {
jobs.value = await fetchJobs();
@@ -157,68 +97,51 @@ async function loadJobs() {
}
}
async function handleSubmit() {
if (form.value.source_type === "telegram" && !form.value.channel.trim()) {
error.value = "Укажите канал Telegram";
return;
}
if (form.value.source_type === "crawl4ai" && !parseLines(form.value.urlsText).length) {
error.value = "Укажите хотя бы один URL";
return;
}
if (form.value.source_type === "viina") {
const urls = parseLines(form.value.urlsText);
const texts = parseLines(form.value.textsText);
if (form.value.input_mode === "urls" && !urls.length) {
error.value = "Укажите URL для VIINA";
return;
}
if (form.value.input_mode === "texts" && !texts.length) {
error.value = "Укажите тексты для VIINA";
return;
}
if (form.value.input_mode === "mixed" && !urls.length && !texts.length) {
error.value = "Укажите URL или тексты";
return;
}
async function loadEntities() {
try {
const [chs, prs] = await Promise.all([fetchChannels(), fetchProfiles()]);
channels.value = chs;
profiles.value = prs;
} catch {
/* selects stay empty */
}
}
async function handleCreatePair() {
if (!pairForm.value.channel_id || !pairForm.value.profile_id) {
error.value = "Выберите канал и профиль";
return;
}
submitting.value = true;
error.value = "";
try {
await createJob({
source_type: form.value.source_type,
source_config: buildSourceConfig(form.value.source_type, form.value),
interval_seconds: form.value.interval_seconds,
channel_id: pairForm.value.channel_id,
profile_id: pairForm.value.profile_id,
limit: pairForm.value.limit,
interval_seconds: pairForm.value.interval_seconds,
is_active: true,
});
form.value.channel = "";
form.value.urlsText = "";
form.value.textsText = "";
pairForm.value.channel_id = null;
pairForm.value.profile_id = null;
await loadJobs();
} catch (err) {
error.value = err instanceof Error ? err.message : "Не удалось создать задание";
error.value = err instanceof Error ? err.message : "Не удалось создать связку";
} finally {
submitting.value = false;
}
}
function openEdit(job: ParseJob) {
if (!isPairJob(job)) {
error.value = "Редактируются только связки канал+профиль. Остальные задания можно удалить.";
return;
}
editingJob.value = job;
editForm.value = {
channel: channelFromConfig(job.source_config),
channel_id: job.channel_id ?? null,
profile_id: job.profile_id ?? null,
limit: limitFromConfig(job.source_config),
urlsText: urlsFromConfig(job.source_config).join("\n"),
textsText: textsFromConfig(job.source_config).join("\n"),
domain_profile:
typeof job.source_config.domain_profile === "string"
? job.source_config.domain_profile
: "generic_news",
input_mode:
job.source_config.input_mode === "texts" || job.source_config.input_mode === "mixed"
? job.source_config.input_mode
: "urls",
extract_mode: job.source_config.extract_mode === "llm" ? "llm" : "heuristic",
interval_seconds: job.interval_seconds,
is_active: job.is_active,
};
@@ -229,19 +152,18 @@ function closeEdit() {
}
async function handleSaveEdit() {
if (!editingJob.value) return;
const sourceType = editingJob.value.source_type;
if (sourceType === "telegram" && !editForm.value.channel.trim()) {
error.value = "Укажите канал Telegram";
if (!editingJob.value || !isPairJob(editingJob.value)) return;
if (!editForm.value.channel_id || !editForm.value.profile_id) {
error.value = "Выберите канал и профиль";
return;
}
submitting.value = true;
error.value = "";
try {
await updateJob(editingJob.value.id, {
source_config: buildSourceConfig(sourceType, editForm.value),
channel_id: editForm.value.channel_id,
profile_id: editForm.value.profile_id,
limit: editForm.value.limit,
interval_seconds: editForm.value.interval_seconds,
is_active: editForm.value.is_active,
});
@@ -268,12 +190,30 @@ async function handleDelete(job: ParseJob) {
}
}
async function handleRetry(jobId: number) {
const STALE_JOB_MS = 15 * 60 * 1000;
function isStaleJob(job: ParseJob): boolean {
if (job.status !== "queued" && job.status !== "running") return false;
if (!job.last_run_at) return job.status === "running";
return Date.now() - new Date(job.last_run_at).getTime() >= STALE_JOB_MS;
}
function canForceRun(job: ParseJob): boolean {
if (job.status !== "queued" && job.status !== "running") return true;
return isStaleJob(job);
}
async function handleForceRun(jobId: number) {
forcingJobId.value = jobId;
error.value = "";
try {
await retryJob(jobId);
await loadJobs();
} catch (err) {
error.value = err instanceof Error ? err.message : "Не удалось перезапустить задание";
error.value =
err instanceof Error ? err.message : "Не удалось запустить запрос постов";
} finally {
forcingJobId.value = null;
}
}
@@ -291,6 +231,7 @@ function formatDate(value: string | null): string {
onMounted(() => {
loadJobs();
loadEntities();
pollTimer = setInterval(loadJobs, 5000);
});
@@ -304,80 +245,43 @@ onUnmounted(() => {
<h2 class="page-heading">Парсеры</h2>
<section class="card">
<h3>Новый парсер</h3>
<h3>Новая связка (канал + профиль)</h3>
<p class="hint">
{{ formHint }}
Дубликаты по <code>source_url</code> не записываются.
Создайте
<router-link to="/channels">канал</router-link>
и
<router-link to="/parser-profiles">профиль</router-link>
(heuristic или llm), затем свяжите их здесь. Flatten в Redis:
heuristic → <code>extract_mode=profile</code>;
llm → <code>extract_mode=llm</code> (batch; listener для llm пока fallback).
</p>
<form class="admin-form" @submit.prevent="handleSubmit">
<form class="admin-form" @submit.prevent="handleCreatePair">
<div class="form-row">
<label>
Тип источника
<select v-model="form.source_type">
<option v-for="opt in SOURCE_TYPES" :key="opt.value" :value="opt.value">
{{ opt.label }}
Канал
<select v-model.number="pairForm.channel_id" required>
<option :value="null" disabled>Выберите…</option>
<option v-for="ch in activeChannels" :key="ch.id" :value="ch.id">
{{ ch.name }} (@{{ ch.channel }})
</option>
</select>
</label>
<template v-if="form.source_type === 'telegram'">
<label>
Канал
<input v-model="form.channel" type="text" placeholder="creamy_caprice" required />
</label>
<label>
Лимит
<input v-model.number="form.limit" type="number" min="1" max="1000" />
</label>
<label>
Извлечение
<select v-model="form.extract_mode">
<option value="heuristic">Эвристики (структурированные посты)</option>
<option value="llm">LLM DeepSeek (неструктурированные)</option>
</select>
</label>
</template>
<template v-else-if="form.source_type === 'crawl4ai'">
<label class="span-2">
URL (по одному в строке)
<textarea v-model="form.urlsText" rows="3" placeholder="https://example.com/news/…" required />
</label>
<label>
Domain profile
<input v-model="form.domain_profile" type="text" placeholder="generic_news" />
</label>
<label>
Извлечение
<select v-model="form.extract_mode">
<option value="heuristic">Эвристики (regex)</option>
<option value="llm">LLM DeepSeek (Crawl4AI)</option>
</select>
</label>
</template>
<template v-else>
<label>
Режим ввода
<select v-model="form.input_mode">
<option value="urls">URLs</option>
<option value="texts">Texts</option>
<option value="mixed">Mixed</option>
</select>
</label>
<label v-if="form.input_mode !== 'texts'" class="span-2">
URL (по одному в строке)
<textarea v-model="form.urlsText" rows="3" placeholder="https://example.com/article…" />
</label>
<label v-if="form.input_mode !== 'urls'" class="span-2">
Тексты (по одному блоку в строке)
<textarea v-model="form.textsText" rows="3" placeholder="Текст статьи…" />
</label>
</template>
<label>
Профиль
<select v-model.number="pairForm.profile_id" required>
<option :value="null" disabled>Выберите…</option>
<option v-for="p in readyProfiles" :key="p.id" :value="p.id">
{{ p.name }} (#{{ p.id }}, {{ p.kind === "llm" ? "llm" : "heuristic" }})
</option>
</select>
</label>
<label>
Лимит batch
<input v-model.number="pairForm.limit" type="number" min="1" max="1000" />
</label>
<label>
Интервал
<select v-model.number="form.interval_seconds">
<select v-model.number="pairForm.interval_seconds">
<option v-for="opt in INTERVAL_OPTIONS" :key="opt.value" :value="opt.value">
{{ opt.label }}
</option>
@@ -385,16 +289,24 @@ onUnmounted(() => {
</label>
</div>
<div class="form-actions">
<button type="submit" class="btn btn-primary" :disabled="submitting">
{{ submitting ? "Создание..." : "Создать парсер" }}
<button
type="submit"
class="btn btn-primary"
:disabled="submitting || !pairForm.channel_id || !pairForm.profile_id"
>
{{ submitting ? "Создание..." : "Создать связку" }}
</button>
</div>
</form>
<p v-if="!activeChannels.length || !readyProfiles.length" class="hint">
Нет активных каналов или готовых профилей —
сначала добавьте их на страницах Каналы / Профили.
</p>
<p v-if="error" class="form-error">{{ error }}</p>
</section>
<section class="card">
<h3>Парсеры <span v-if="loading" class="muted">(загрузка...)</span></h3>
<h3>Связки <span v-if="loading" class="muted">(загрузка...)</span></h3>
<div class="table-wrap">
<table class="admin-table">
<thead>
@@ -426,19 +338,26 @@ onUnmounted(() => {
<td>{{ formatDate(job.last_run_at) }}</td>
<td class="error-cell">{{ job.last_error ?? "—" }}</td>
<td class="actions-cell">
<button class="btn btn-sm" type="button" @click="openEdit(job)">Изменить</button>
<button
v-if="job.status === 'failed' || job.status === 'error'"
v-if="isPairJob(job)"
class="btn btn-sm"
type="button"
@click="handleRetry(job.id)"
@click="openEdit(job)"
>
Повторить
Изменить
</button>
<button
class="btn btn-sm"
type="button"
:disabled="!canForceRun(job) || forcingJobId === job.id"
@click="handleForceRun(job.id)"
>
{{ forcingJobId === job.id ? "Запуск..." : "Запустить" }}
</button>
<button
class="btn btn-sm btn-danger"
type="button"
:disabled="job.status === 'running'"
:disabled="job.status === 'running' && !isStaleJob(job)"
@click="handleDelete(job)"
>
Удалить
@@ -452,61 +371,29 @@ onUnmounted(() => {
<div v-if="editingJob" class="modal-backdrop" @click.self="closeEdit">
<div class="modal card">
<h3>Редактировать парсер #{{ editingJob.id }} ({{ editingJob.source_type }})</h3>
<h3>Редактировать связку #{{ editingJob.id }}</h3>
<form class="admin-form" @submit.prevent="handleSaveEdit">
<div class="form-row">
<template v-if="editingJob.source_type === 'telegram'">
<label>
Канал
<input v-model="editForm.channel" type="text" required />
</label>
<label>
Лимит
<input v-model.number="editForm.limit" type="number" min="1" max="1000" />
</label>
<label>
Извлечение
<select v-model="editForm.extract_mode">
<option value="heuristic">Эвристики</option>
<option value="llm">LLM DeepSeek</option>
</select>
</label>
</template>
<template v-else-if="editingJob.source_type === 'crawl4ai'">
<label class="span-2">
URL (по одному в строке)
<textarea v-model="editForm.urlsText" rows="3" required />
</label>
<label>
Domain profile
<input v-model="editForm.domain_profile" type="text" />
</label>
<label>
Извлечение
<select v-model="editForm.extract_mode">
<option value="heuristic">Эвристики</option>
<option value="llm">LLM DeepSeek</option>
</select>
</label>
</template>
<template v-else>
<label>
Режим ввода
<select v-model="editForm.input_mode">
<option value="urls">URLs</option>
<option value="texts">Texts</option>
<option value="mixed">Mixed</option>
</select>
</label>
<label v-if="editForm.input_mode !== 'texts'" class="span-2">
URL
<textarea v-model="editForm.urlsText" rows="3" />
</label>
<label v-if="editForm.input_mode !== 'urls'" class="span-2">
Тексты
<textarea v-model="editForm.textsText" rows="3" />
</label>
</template>
<label>
Канал
<select v-model.number="editForm.channel_id" required>
<option v-for="ch in editChannelOptions" :key="ch.id" :value="ch.id">
{{ ch.name }} (@{{ ch.channel }}){{ ch.is_active ? "" : " — выкл." }}
</option>
</select>
</label>
<label>
Профиль
<select v-model.number="editForm.profile_id" required>
<option v-for="p in editProfileOptions" :key="p.id" :value="p.id">
{{ p.name }} (#{{ p.id }})
</option>
</select>
</label>
<label>
Лимит batch
<input v-model.number="editForm.limit" type="number" min="1" max="1000" />
</label>
<label>
Интервал
<select v-model.number="editForm.interval_seconds">
@@ -540,19 +427,6 @@ onUnmounted(() => {
line-height: 1.5;
}
.form-row .span-2 {
grid-column: span 2;
}
.form-row textarea {
width: 100%;
font: inherit;
padding: 0.4rem 0.5rem;
border: 1px solid #d1d5db;
border-radius: 4px;
resize: vertical;
}
.source-cell {
max-width: 280px;
overflow: hidden;
@@ -568,11 +442,6 @@ onUnmounted(() => {
margin-right: 0.25rem;
}
.btn-danger {
color: #b91c1c;
border-color: #fecaca;
}
.modal-backdrop {
position: fixed;
inset: 0;
@@ -586,8 +455,6 @@ onUnmounted(() => {
.modal {
width: min(560px, 92vw);
margin: 0;
max-height: 90vh;
overflow: auto;
}
.checkbox-row {
@@ -596,4 +463,13 @@ onUnmounted(() => {
gap: 0.5rem;
margin-top: 1.5rem;
}
.error-cell {
max-width: 200px;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
font-size: 0.8rem;
color: #b91c1c;
}
</style>
@@ -0,0 +1,278 @@
<script setup lang="ts">
import { computed, onMounted, ref } from "vue";
import { fetchVpnSettings, updateVpnSettings } from "../api/admin";
import type { VpnMode, VpnSettings } from "../types/admin";
const loading = ref(true);
const submitting = ref(false);
const error = ref("");
const success = ref("");
const settings = ref<VpnSettings | null>(null);
const form = ref({
enabled: false,
mode: "subscription" as VpnMode,
subscription_url: "",
clear_subscription_url: false,
subscription_interval_seconds: 3600,
host: "",
port: 1080 as number | null,
username: "",
password: "",
clear_password: false,
proxied_source_types: [] as string[],
});
const sourceLabels: Record<string, string> = {
telegram: "Telegram",
crawl4ai: "Crawl4AI (web)",
viina: "VIINA (nlp)",
};
const isSubscription = computed(() => form.value.mode === "subscription");
const isDirectProxy = computed(
() => form.value.mode === "socks5" || form.value.mode === "http",
);
function applySettings(data: VpnSettings) {
settings.value = data;
form.value = {
enabled: data.enabled,
mode: data.mode,
subscription_url: "",
clear_subscription_url: false,
subscription_interval_seconds: data.subscription_interval_seconds,
host: data.host ?? "",
port: data.port ?? 1080,
username: data.username ?? "",
password: "",
clear_password: false,
proxied_source_types: [...data.proxied_source_types],
};
}
async function load() {
loading.value = true;
error.value = "";
success.value = "";
try {
applySettings(await fetchVpnSettings());
} catch (err) {
error.value = err instanceof Error ? err.message : "Не удалось загрузить VPN";
} finally {
loading.value = false;
}
}
function toggleSource(sourceType: string, checked: boolean) {
const set = new Set(form.value.proxied_source_types);
if (checked) set.add(sourceType);
else set.delete(sourceType);
form.value.proxied_source_types = [...set];
}
async function handleSave() {
submitting.value = true;
error.value = "";
success.value = "";
try {
const payload: Parameters<typeof updateVpnSettings>[0] = {
enabled: form.value.enabled,
mode: form.value.mode,
subscription_interval_seconds: form.value.subscription_interval_seconds,
host: form.value.host.trim() || null,
port: form.value.port,
username: form.value.username.trim() || null,
proxied_source_types: form.value.proxied_source_types,
};
if (form.value.clear_subscription_url) {
payload.clear_subscription_url = true;
} else if (form.value.subscription_url.trim()) {
payload.subscription_url = form.value.subscription_url.trim();
}
if (form.value.clear_password) {
payload.clear_password = true;
} else if (form.value.password) {
payload.password = form.value.password;
}
applySettings(await updateVpnSettings(payload));
success.value = "Настройки VPN сохранены";
} catch (err) {
error.value = err instanceof Error ? err.message : "Не удалось сохранить";
} finally {
submitting.value = false;
}
}
onMounted(load);
</script>
<template>
<div class="page">
<h2 class="page-heading">VPN</h2>
<p class="intro">
Проксирование трафика выбранных источников парсинга. Режим
<strong>subscription</strong> поднимает SOCKS через сервис
<code>cp-vpn</code> (mihomo). Режимы SOCKS5/HTTP задают адрес напрямую для
воркеров. Без включённого VPN остаётся fallback из
<code>TELEGRAM_PROXY_*</code> в <code>.env</code>.
</p>
<section class="card">
<h3>Настройки <span v-if="loading" class="muted">(загрузка...)</span></h3>
<form v-if="settings" class="admin-form" @submit.prevent="handleSave">
<label class="checkbox-row">
<input v-model="form.enabled" type="checkbox" />
<span>Прокси включён</span>
</label>
<div class="form-row vpn-row">
<label>
Способ проксирования
<select v-model="form.mode">
<option value="subscription">Subscription URL</option>
<option value="socks5">SOCKS5</option>
<option value="http">HTTP</option>
</select>
</label>
</div>
<template v-if="isSubscription">
<label>
Адрес подписки (subscription URL)
<input
v-model="form.subscription_url"
type="url"
placeholder="https://sub.example.com/…"
autocomplete="off"
/>
</label>
<p v-if="settings.subscription_url_set" class="hint">
Сохранён:
<code>{{ settings.subscription_url_hint }}</code>
— оставьте поле пустым, чтобы не менять.
</p>
<label class="checkbox-row">
<input v-model="form.clear_subscription_url" type="checkbox" />
<span>Очистить сохранённый URL</span>
</label>
<label>
Интервал обновления подписки (сек)
<input
v-model.number="form.subscription_interval_seconds"
type="number"
min="60"
max="86400"
/>
</label>
</template>
<template v-if="isDirectProxy">
<div class="form-row">
<label>
Host
<input v-model="form.host" type="text" placeholder="127.0.0.1" />
</label>
<label>
Port
<input v-model.number="form.port" type="number" min="1" max="65535" />
</label>
<label>
Username
<input v-model="form.username" type="text" autocomplete="off" />
</label>
<label>
Password
<input
v-model="form.password"
type="password"
autocomplete="new-password"
:placeholder="settings.password_set ? '•••• (оставьте пустым)' : ''"
/>
</label>
</div>
<label v-if="settings.password_set" class="checkbox-row">
<input v-model="form.clear_password" type="checkbox" />
<span>Очистить пароль</span>
</label>
</template>
<fieldset class="sources">
<legend>Источники через прокси</legend>
<label
v-for="st in settings.available_source_types"
:key="st"
class="checkbox-row"
>
<input
type="checkbox"
:checked="form.proxied_source_types.includes(st)"
@change="
toggleSource(st, ($event.target as HTMLInputElement).checked)
"
/>
<span>{{ sourceLabels[st] || st }} <code>({{ st }})</code></span>
</label>
</fieldset>
<div class="form-actions">
<button type="submit" class="btn btn-primary" :disabled="submitting || loading">
{{ submitting ? "Сохранение..." : "Сохранить" }}
</button>
</div>
<p v-if="error" class="form-error">{{ error }}</p>
<p v-if="success" class="form-success">{{ success }}</p>
</form>
<p v-else-if="error" class="form-error">{{ error }}</p>
</section>
</div>
</template>
<style scoped>
.intro {
margin: 0 0 1rem;
font-size: 0.9rem;
color: #555;
line-height: 1.5;
}
.checkbox-row {
display: flex;
align-items: center;
gap: 0.5rem;
margin: 0.5rem 0 1rem;
}
.vpn-row {
margin-bottom: 0.75rem;
}
.hint {
margin: -0.5rem 0 0.75rem;
font-size: 0.85rem;
color: #666;
}
.sources {
border: 1px solid #e5e5e5;
border-radius: 6px;
padding: 0.75rem 1rem 0.25rem;
margin: 0.5rem 0 1rem;
}
.sources legend {
padding: 0 0.35rem;
font-size: 0.9rem;
color: #444;
}
.form-success {
margin: 0.75rem 0 0;
color: #1a7f37;
font-size: 0.9rem;
}
</style>
+35 -4
View File
@@ -54,14 +54,42 @@ async def run(job_id, source_config, *, ctx) -> tuple[list[dict], str | None]
- `source_type` события = тип адаптера;
- `source_config` валидируется схемами из `contracts/sources.py`.
### Режимы извлечения (Telegram)
Цель всегда одна: фиксированные поля `IngestEventItem` / Event в ЦА (`title`, `description`, `locality`, `coords`→lat/lng, `event_date`, `topic`, …). Кастомные пользовательские таблицы — roadmap.
| `extract_mode` | Что делает | Откуда в UI |
|----------------|------------|-------------|
| `heuristic` | Legacy regex-парсер постов | raw job API (не пара) |
| `llm` | DeepSeek **на каждый** пост (batch) | профиль `kind=llm` → pair flatten |
| `profile` | Статичные правила `HeuristicProfile` | профиль `kind=heuristic` → pair flatten |
### LLM-режим (`extract_mode: llm`)
Для неструктурированных Telegram-постов и Crawl4AI:
- ключ `DEEPSEEK_API_KEY` в `.env` (воркеры `cp-workers` / `cp-workers-web`);
- Telegram: текст поста → DeepSeek JSON → `IngestEventItem`;
- ключ `DEEPSEEK_API_KEY` в `.env` (воркеры `cp-workers` / `cp-workers-web`; preview в `ca-api`);
- Telegram batch: текст → DeepSeek JSON → один или несколько `IngestEventItem` (`workers/llm_extract.py`);
- опционально `extract_schema`, `instruction`, `required_fields` (post-extract gate поверх `is_event`);
- **`multi_event`** (opt-in в `llm_profile` / `TelegramSourceConfig`): модель возвращает массив `events`; каждое событие — отдельный ingest. URL: один event → `post.url`; несколько → `post.url#e1`, `#e2`, … (дедуп ЦА по `source_url`);
- heuristic / profile по-прежнему **1 пост → 1 событие**;
- Crawl4AI: страница → `LLMExtractionStrategy` (DeepSeek) с fallback на тот же DeepSeek по markdown;
- в UI «Парсеры»: поле **Извлечение** = LLM DeepSeek.
- **listener:** для `llm` пока fallback на heuristic (LLM — batch-only by design).
Reusable LLM-профиль в ЦА: `ParserProfile.kind=llm` + JSON `llm_profile` (`contracts/llm_profile.py`). При enqueue flatten → `extract_mode=llm` + schema/instruction/required_fields/`multi_event`.
### Profile-режим (`extract_mode: profile`)
Статичный парсер из конструктора ЦА:
- `heuristic_profile` в `source_config` (схема `contracts/heuristic_profile.py`);
- интерпретатор: `workers/heuristic_profile.py` (те же правила, что preview в ЦА);
- `required_fields`: пост без заполненных обязательных полей не ингестится (batch + listener);
- генерация правил — один раз в админке (`/admin/parser-profiles/generate`); runtime **без** LLM.
Reusable heuristic-профиль: `ParserProfile.kind=heuristic` + `heuristic_profile`. Pair на «Парсеры» → flatten `extract_mode=profile`.
UI «Профили»: выбор `kind` (heuristic | llm). Связка канал+профиль на «Парсеры».
```mermaid
flowchart LR
@@ -118,7 +146,10 @@ Legacy-ключ `cp:jobs` по-прежнему дренируется telegram-
## Telegram real-time
`TelegramListener` без изменений: подписки из ЦА, ingest с `listener: true` (статус `ParseJob` не трогается).
`TelegramListener`: подписки из ЦА, ingest с `listener: true` (статус `ParseJob` не трогается).
- `extract_mode=profile` — те же heuristic-правила, что batch;
- `extract_mode=llm` — **не** вызывает DeepSeek; fallback на legacy heuristic (LLM только в batch).
## Переменные окружения
+31
View File
@@ -0,0 +1,31 @@
FROM python:3.12-alpine
ARG MIHOMO_VERSION=1.19.3
ARG TARGETARCH
RUN apk add --no-cache ca-certificates curl gzip \
&& arch="${TARGETARCH:-amd64}" \
&& case "$arch" in \
amd64|x86_64) mihomo_arch=amd64 ;; \
arm64|aarch64) mihomo_arch=arm64 ;; \
*) mihomo_arch=amd64 ;; \
esac \
&& curl -fsSL \
"https://github.com/MetaCubeX/mihomo/releases/download/v${MIHOMO_VERSION}/mihomo-linux-${mihomo_arch}-compatible-v${MIHOMO_VERSION}.gz" \
-o /tmp/mihomo.gz \
&& gunzip -c /tmp/mihomo.gz > /usr/local/bin/mihomo \
&& chmod +x /usr/local/bin/mihomo \
&& rm -f /tmp/mihomo.gz \
&& mkdir -p /etc/mihomo/providers
WORKDIR /app
COPY controller.py /app/controller.py
ENV MIHOMO_BIN=/usr/local/bin/mihomo \
MIHOMO_CONFIG_DIR=/etc/mihomo \
VPN_SOCKS_PORT=1080 \
VPN_POLL_SECONDS=15
EXPOSE 1080
CMD ["python", "-u", "/app/controller.py"]
+204
View File
@@ -0,0 +1,204 @@
"""Poll CA /internal/vpn and keep mihomo config in sync for subscription mode."""
from __future__ import annotations
import json
import logging
import os
import signal
import subprocess
import sys
import time
from pathlib import Path
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen
logging.basicConfig(
level=logging.INFO,
format="%(asctime)s %(levelname)s %(message)s",
)
logger = logging.getLogger("cp-vpn")
CA_API_URL = os.getenv("CA_API_URL", "http://ca-api:8000").rstrip("/")
INTERNAL_TOKEN = os.getenv("INTERNAL_TOKEN", "dev-internal-token")
POLL_SECONDS = int(os.getenv("VPN_POLL_SECONDS", "15"))
CONFIG_DIR = Path(os.getenv("MIHOMO_CONFIG_DIR", "/etc/mihomo"))
CONFIG_PATH = CONFIG_DIR / "config.yaml"
MIHOMO_BIN = os.getenv("MIHOMO_BIN", "/usr/local/bin/mihomo")
MIXED_PORT = int(os.getenv("VPN_SOCKS_PORT", "1080"))
CONTROLLER = "127.0.0.1:9090"
_mihomo_proc: subprocess.Popen | None = None
_last_fingerprint: str | None = None
def fetch_vpn() -> dict:
req = Request(
f"{CA_API_URL}/internal/vpn",
headers={"X-Internal-Token": INTERNAL_TOKEN},
method="GET",
)
with urlopen(req, timeout=15) as resp:
return json.loads(resp.read().decode("utf-8"))
def fingerprint(cfg: dict) -> str:
return json.dumps(
{
"enabled": cfg.get("enabled"),
"mode": cfg.get("mode"),
"subscription_url": cfg.get("subscription_url"),
"subscription_interval_seconds": cfg.get("subscription_interval_seconds"),
},
sort_keys=True,
)
def write_direct_config() -> None:
CONFIG_DIR.mkdir(parents=True, exist_ok=True)
(CONFIG_DIR / "providers").mkdir(parents=True, exist_ok=True)
CONFIG_PATH.write_text(
f"""# managed by cp-vpn controller — DIRECT (subscription inactive)
mixed-port: {MIXED_PORT}
allow-lan: true
bind-address: "*"
mode: direct
log-level: warning
external-controller: {CONTROLLER}
""",
encoding="utf-8",
)
def write_subscription_config(url: str, interval: int) -> None:
CONFIG_DIR.mkdir(parents=True, exist_ok=True)
(CONFIG_DIR / "providers").mkdir(parents=True, exist_ok=True)
# YAML: quote URL; escape double quotes in URL if any
safe_url = url.replace("\\", "\\\\").replace('"', '\\"')
CONFIG_PATH.write_text(
f"""# managed by cp-vpn controller — subscription
mixed-port: {MIXED_PORT}
allow-lan: true
bind-address: "*"
mode: rule
log-level: info
external-controller: {CONTROLLER}
proxy-providers:
sub:
type: http
url: "{safe_url}"
interval: {interval}
path: ./providers/sub.yaml
health-check:
enable: true
url: https://www.gstatic.com/generate_204
interval: 600
proxy-groups:
- name: PROXY
type: select
use:
- sub
rules:
- MATCH,PROXY
""",
encoding="utf-8",
)
def start_mihomo() -> None:
global _mihomo_proc
if _mihomo_proc is not None and _mihomo_proc.poll() is None:
return
logger.info("Starting mihomo (%s)", MIHOMO_BIN)
_mihomo_proc = subprocess.Popen(
[MIHOMO_BIN, "-d", str(CONFIG_DIR)],
stdout=sys.stdout,
stderr=sys.stderr,
)
def stop_mihomo() -> None:
global _mihomo_proc
if _mihomo_proc is None:
return
if _mihomo_proc.poll() is None:
logger.info("Stopping mihomo")
_mihomo_proc.send_signal(signal.SIGTERM)
try:
_mihomo_proc.wait(timeout=10)
except subprocess.TimeoutExpired:
_mihomo_proc.kill()
_mihomo_proc = None
def reload_mihomo() -> None:
"""Force reload via external controller; fall back to process restart."""
body = json.dumps({"path": str(CONFIG_PATH)}).encode("utf-8")
req = Request(
f"http://{CONTROLLER}/configs?force=true",
data=body,
headers={"Content-Type": "application/json"},
method="PUT",
)
try:
with urlopen(req, timeout=10) as resp:
resp.read()
logger.info("Mihomo config reloaded")
return
except (HTTPError, URLError, OSError) as exc:
logger.warning("Reload via API failed (%s); restarting mihomo", exc)
stop_mihomo()
start_mihomo()
def apply_config(cfg: dict) -> None:
global _last_fingerprint
fp = fingerprint(cfg)
if fp == _last_fingerprint and _mihomo_proc is not None and _mihomo_proc.poll() is None:
return
enabled = bool(cfg.get("enabled"))
mode = (cfg.get("mode") or "").strip()
url = (cfg.get("subscription_url") or "").strip()
interval = int(cfg.get("subscription_interval_seconds") or 3600)
if enabled and mode == "subscription" and url:
logger.info("Applying subscription config (interval=%ss)", interval)
write_subscription_config(url, interval)
else:
logger.info("Applying DIRECT config (enabled=%s mode=%s)", enabled, mode)
write_direct_config()
already_running = _mihomo_proc is not None and _mihomo_proc.poll() is None
start_mihomo()
if already_running or _last_fingerprint is not None:
time.sleep(1)
reload_mihomo()
_last_fingerprint = fp
def main() -> None:
global _mihomo_proc
logger.info(
"cp-vpn controller started (poll=%ss api=%s)",
POLL_SECONDS,
CA_API_URL,
)
write_direct_config()
start_mihomo()
while True:
try:
cfg = fetch_vpn()
apply_config(cfg)
except Exception:
logger.exception("Failed to sync VPN config")
if _mihomo_proc is not None and _mihomo_proc.poll() is not None:
logger.warning("mihomo exited with code %s; restarting", _mihomo_proc.returncode)
_mihomo_proc = None
start_mihomo()
time.sleep(POLL_SECONDS)
if __name__ == "__main__":
main()
+22 -1
View File
@@ -179,10 +179,31 @@ async def run_with_listener() -> None:
await worker_loop(ctx=WorkerContext())
return
from workers.sources.telegram_client import TelegramAuthError, TelegramConfigError
from workers.sources.telegram_listener import TelegramListener
from workers.sources.telegram_session import close_shared_client, get_shared_client
client = await get_shared_client()
while True:
try:
client = await get_shared_client()
break
except (TelegramAuthError, TelegramConfigError, ConnectionError, OSError) as exc:
logger.error(
"Telegram connect failed (%s); retry in 60s (batch without TG)",
exc,
)
try:
await asyncio.wait_for(
worker_loop(ctx=WorkerContext(tg_client=None)),
timeout=60,
)
except asyncio.TimeoutError:
pass
continue
except Exception:
logger.exception("Unexpected Telegram connect error; retry in 60s")
await asyncio.sleep(60)
listener = TelegramListener(client)
ctx = WorkerContext(tg_client=client)
worker_task = asyncio.create_task(worker_loop(ctx=ctx))
@@ -57,7 +57,25 @@ class Crawl4AIAdapter:
events: list[dict] = []
errors: list[str] = []
async with AsyncWebCrawler(verbose=False) as crawler:
crawler_kwargs: dict[str, Any] = {"verbose": False}
try:
from workers.sources.vpn_config import http_proxy_url
proxy = http_proxy_url("crawl4ai")
if proxy:
try:
from crawl4ai import BrowserConfig # type: ignore
crawler_kwargs["config"] = BrowserConfig(proxy=proxy)
logger.info("Crawl4AI via proxy %s", proxy.split("@")[-1])
except Exception:
logger.exception(
"Could not apply VPN proxy to Crawl4AI; continuing without"
)
except Exception:
logger.exception("VPN config lookup failed for crawl4ai")
async with AsyncWebCrawler(**crawler_kwargs) as crawler:
for url in cfg.urls:
try:
if cfg.extract_mode == "llm":
@@ -7,12 +7,15 @@ import logging
from contracts.sources import TelegramSourceConfig
from workers.adapters.base import WorkerContext
from workers.converter import event_record_to_ingest
from workers.heuristic_profile import extract_with_profile
from workers.llm_extract import (
DEFAULT_EXTRACT_SCHEMA,
DEFAULT_INSTRUCTION,
extract_event_fields,
event_source_url,
extract_event_list,
fields_to_ingest,
llm_enabled,
match_llm_required,
)
from workers.parsers.telegram_events import parse_event_posts
from workers.sources.telegram_client import (
@@ -57,43 +60,87 @@ class TelegramAdapter:
return [], "extract_mode=llm requires DEEPSEEK_API_KEY in worker env"
return await _extract_posts_with_llm(posts, cfg)
if cfg.extract_mode == "profile":
return _extract_posts_with_profile(posts, cfg)
records = parse_event_posts(posts)
events = [event_record_to_ingest(r) for r in records]
return events, None
def _extract_posts_with_profile(posts, cfg: TelegramSourceConfig) -> tuple[list[dict], str | None]:
profile = cfg.heuristic_profile or {}
events: list[dict] = []
for post in posts:
text = (post.text or "").strip()
if not text:
continue
event = extract_with_profile(
text,
profile,
source_url=post.url,
source_type="telegram",
extra_metadata={
"channel": post.channel,
"message_id": post.id,
"post_date": post.date.isoformat() if post.date else None,
},
)
if event is None:
logger.debug("Profile skip %s (required fields missing)", post.url)
continue
events.append(event)
return events, None
async def _extract_posts_with_llm(posts, cfg: TelegramSourceConfig) -> tuple[list[dict], str | None]:
schema = cfg.extract_schema or DEFAULT_EXTRACT_SCHEMA
instruction = cfg.instruction or DEFAULT_INSTRUCTION
multi_event = bool(cfg.multi_event)
events: list[dict] = []
errors: list[str] = []
required = list(cfg.required_fields or [])
for post in posts:
text = (post.text or "").strip()
if not text:
continue
try:
fields = await extract_event_fields(
field_list = await extract_event_list(
text,
extract_schema=schema,
instruction=instruction,
multi_event=multi_event,
)
if not fields.get("is_event", True):
if not field_list:
continue
events.append(
fields_to_ingest(
source_type="telegram",
source_url=post.url,
raw_text=text,
fields=fields,
domain_profile="telegram_llm",
extra_metadata={
"channel": post.channel,
"message_id": post.id,
"post_date": post.date.isoformat() if post.date else None,
},
total = len(field_list)
for idx, fields in enumerate(field_list):
if not match_llm_required(fields, required):
logger.debug(
"LLM skip %s event %s/%s (required fields missing)",
post.url,
idx + 1,
total,
)
continue
events.append(
fields_to_ingest(
source_type="telegram",
source_url=event_source_url(post.url, idx, total),
raw_text=text,
fields=fields,
domain_profile="telegram_llm",
extra_metadata={
"channel": post.channel,
"message_id": post.id,
"post_date": post.date.isoformat() if post.date else None,
"event_index": idx + 1,
"event_count": total,
"multi_event": multi_event and total > 1,
},
)
)
)
except Exception as exc:
logger.exception("LLM extract failed for %s", post.url)
errors.append(f"{post.url}: {exc}")
@@ -70,7 +70,14 @@ class ViinaAdapter:
async def _fetch_article_text(url: str) -> str:
async with httpx.AsyncClient(timeout=60.0, follow_redirects=True) as client:
from workers.sources.vpn_config import http_proxy_url
proxy = http_proxy_url("viina")
kwargs: dict = {"timeout": 60.0, "follow_redirects": True}
if proxy:
kwargs["proxy"] = proxy
logger.info("VIINA fetch via proxy %s", proxy.split("@")[-1])
async with httpx.AsyncClient(**kwargs) as client:
response = await client.get(
url,
headers={"User-Agent": "MapMil-CP-Viina/1.0"},
@@ -0,0 +1,75 @@
"""Apply HeuristicProfile to Telegram posts → IngestEventItem-shaped dicts."""
from __future__ import annotations
from typing import Any
from contracts.heuristic_profile import HeuristicProfile, match_profile
from workers.llm_extract import parse_coords, parse_date
def profile_fields_to_ingest(
*,
source_type: str,
source_url: str,
raw_text: str,
fields: dict[str, str],
extra_metadata: dict | None = None,
) -> dict:
lat, lng = parse_coords(str(fields.get("coords") or ""))
description = str(fields.get("description") or raw_text)[:8000]
locality = str(fields.get("locality") or "")
region = str(fields.get("region") or locality or "") or None
if fields.get("title"):
title = str(fields["title"])
elif locality:
title = locality
elif description:
title = description.splitlines()[0][:120]
else:
title = ""
topic = str(fields.get("topic") or "telegram_profile")
event_date = parse_date(str(fields.get("event_date") or ""))
meta: dict[str, Any] = {
"extract_mode": "profile",
"extracted": dict(fields),
}
if extra_metadata:
meta.update(extra_metadata)
return {
"source_type": source_type,
"source_url": source_url,
"raw_text": raw_text[:20000],
"title": title[:255],
"description": description,
"locality": locality,
"latitude": lat,
"longitude": lng,
"event_date": event_date.isoformat() if event_date else None,
"region": region,
"topic": topic,
"tags": [source_type, "profile"],
"metadata": meta,
}
def extract_with_profile(
text: str,
profile: HeuristicProfile | dict[str, Any],
*,
source_url: str,
source_type: str = "telegram",
extra_metadata: dict | None = None,
) -> dict | None:
matched, fields, _missing = match_profile(text, profile)
if not matched:
return None
return profile_fields_to_ingest(
source_type=source_type,
source_url=source_url,
raw_text=text,
fields=fields,
extra_metadata=extra_metadata,
)
+128 -48
View File
@@ -5,30 +5,31 @@ from __future__ import annotations
import json
import logging
import os
import re
from datetime import datetime, timezone
from typing import Any
import httpx
from contracts.heuristic_profile import parse_coords
from contracts.llm_profile import (
DEFAULT_EXTRACT_SCHEMA,
DEFAULT_INSTRUCTION,
match_llm_required,
)
logger = logging.getLogger("cp-worker.llm")
COORDS_RE = re.compile(r"(-?\d{1,3}\.\d+)\s*,\s*(-?\d{1,3}\.\d+)")
DEFAULT_EXTRACT_SCHEMA: dict[str, str] = {
"title": "string — short event title",
"locality": "string — place / settlement name",
"event_date": "string — date as DD.MM.YYYY or YYYY-MM-DD if known",
"description": "string — concise event summary",
"coords": "string — latitude, longitude if present else empty",
"topic": "string — short topic tag",
}
DEFAULT_INSTRUCTION = (
"Extract structured military/news event fields from the text. "
"If the text is not an event, return is_event=false. "
"Respond with a single JSON object only."
)
# Re-export for adapters that import from this module
__all__ = [
"DEFAULT_EXTRACT_SCHEMA",
"DEFAULT_INSTRUCTION",
"event_source_url",
"extract_event_fields",
"extract_event_list",
"fields_to_ingest",
"llm_enabled",
"match_llm_required",
]
def llm_enabled() -> bool:
@@ -43,29 +44,57 @@ def llm_settings() -> dict[str, str]:
}
async def extract_event_fields(
text: str,
def event_source_url(post_url: str, index: int, total: int) -> str:
"""Stable per-event URL for CA dedup. Single event keeps bare post URL."""
base = (post_url or "").strip()
if total <= 1:
return base
# index is 0-based; fragment uses 1-based #eN
return f"{base}#e{index + 1}"
def _normalize_fields(raw: dict[str, Any], schema: dict[str, str]) -> dict[str, str]:
return {key: str(raw.get(key) or "").strip() for key in schema}
def _parse_llm_content(
content: str,
schema: dict[str, str],
*,
extract_schema: dict[str, str] | None = None,
instruction: str | None = None,
) -> dict[str, Any]:
"""Ask DeepSeek to fill schema fields from free text. Returns dict (+ is_event)."""
multi_event: bool,
) -> tuple[bool, list[dict[str, str]]]:
"""Return (is_event, list of field dicts)."""
parsed = json.loads(content)
if not isinstance(parsed, dict):
return False, []
is_event = bool(parsed.get("is_event", True))
if not is_event:
return False, []
if multi_event and isinstance(parsed.get("events"), list):
events: list[dict[str, str]] = []
for item in parsed["events"]:
if isinstance(item, dict):
events.append(_normalize_fields(item, schema))
return True, events
# Single-event fallback (legacy fields / top-level keys)
fields_raw = parsed.get("fields") if isinstance(parsed.get("fields"), dict) else parsed
if not isinstance(fields_raw, dict):
fields_raw = {}
# Drop non-schema keys that confuse normalize when falling back to top-level
cleaned = {k: fields_raw.get(k) for k in schema}
return True, [_normalize_fields(cleaned, schema)]
async def _call_deepseek(user_prompt: str) -> str:
settings = llm_settings()
if not settings["api_key"]:
raise RuntimeError(
"DEEPSEEK_API_KEY is not set. Add it to .env for LLM extract_mode."
)
schema = extract_schema or DEFAULT_EXTRACT_SCHEMA
instr = instruction or DEFAULT_INSTRUCTION
schema_lines = "\n".join(f"- {k}: {v}" for k, v in schema.items())
user_prompt = (
f"{instr}\n\n"
f"Fields to extract:\n{schema_lines}\n\n"
'Return JSON: {"is_event": true|false, "fields": {<field>: <string>}}\n\n'
f"Text:\n{text[:12000]}"
)
payload = {
"model": settings["model"],
"messages": [
@@ -95,14 +124,72 @@ async def extract_event_fields(
response.raise_for_status()
data = response.json()
content = data["choices"][0]["message"]["content"]
parsed = json.loads(content)
fields = parsed.get("fields") if isinstance(parsed.get("fields"), dict) else parsed
if not isinstance(fields, dict):
fields = {}
# Normalize to strings for known keys
result = {key: str(fields.get(key) or "").strip() for key in schema}
result["is_event"] = bool(parsed.get("is_event", True))
return data["choices"][0]["message"]["content"]
async def extract_event_list(
text: str,
*,
extract_schema: dict[str, str] | None = None,
instruction: str | None = None,
multi_event: bool = False,
) -> list[dict[str, Any]]:
"""Extract zero or more event field dicts from free text.
Each dict has schema keys as strings. Empty list if not an event / no items.
"""
schema = extract_schema or DEFAULT_EXTRACT_SCHEMA
instr = instruction or DEFAULT_INSTRUCTION
schema_lines = "\n".join(f"- {k}: {v}" for k, v in schema.items())
if multi_event:
user_prompt = (
f"{instr}\n\n"
"If the text describes multiple distinct events (different places, "
"coords, or dates), return one object per event in \"events\".\n"
f"Fields per event:\n{schema_lines}\n\n"
'Return JSON: {"is_event": true|false, "events": [{<field>: <string>}, ...]}\n'
"If there is no event, return is_event=false and events=[].\n\n"
f"Text:\n{text[:12000]}"
)
else:
user_prompt = (
f"{instr}\n\n"
f"Fields to extract:\n{schema_lines}\n\n"
'Return JSON: {"is_event": true|false, "fields": {<field>: <string>}}\n\n'
f"Text:\n{text[:12000]}"
)
content = await _call_deepseek(user_prompt)
is_event, events = _parse_llm_content(content, schema, multi_event=multi_event)
if not is_event:
return []
return events
async def extract_event_fields(
text: str,
*,
extract_schema: dict[str, str] | None = None,
instruction: str | None = None,
) -> dict[str, Any]:
"""Ask DeepSeek to fill schema fields from free text. Returns dict (+ is_event).
Single-event API kept for crawl4ai and legacy callers.
"""
events = await extract_event_list(
text,
extract_schema=extract_schema,
instruction=instruction,
multi_event=False,
)
if not events:
schema = extract_schema or DEFAULT_EXTRACT_SCHEMA
empty = {key: "" for key in schema}
empty["is_event"] = False
return empty
result = dict(events[0])
result["is_event"] = True
return result
@@ -146,13 +233,6 @@ def fields_to_ingest(
}
def parse_coords(raw: str) -> tuple[float | None, float | None]:
match = COORDS_RE.search(raw or "")
if not match:
return None, None
return float(match.group(1)), float(match.group(2))
def parse_date(raw: str) -> datetime | None:
if not raw:
return None
@@ -9,6 +9,7 @@ import httpx
from telethon import TelegramClient, events
from workers.converter import event_record_to_ingest
from workers.heuristic_profile import extract_with_profile
from workers.parsers.telegram_events import parse_event_post
from workers.sources.telegram_client import message_to_post, normalize_channel
@@ -22,11 +23,12 @@ REFRESH_SECONDS = int(os.getenv("TELEGRAM_LISTENER_REFRESH_SECONDS", "60"))
class TelegramListener:
def __init__(self, client: TelegramClient) -> None:
self.client = client
self._channels: dict[str, int] = {}
# channel -> {job_id, source_config}
self._channels: dict[str, dict[str, Any]] = {}
self._chat_ids: set[int] = set()
self._handlers_registered = False
async def fetch_subscriptions(self) -> dict[str, int]:
async def fetch_subscriptions(self) -> dict[str, dict[str, Any]]:
async with httpx.AsyncClient(timeout=30.0) as http:
response = await http.get(
f"{CA_API_URL}/internal/listener/subscriptions",
@@ -35,7 +37,7 @@ class TelegramListener:
response.raise_for_status()
data = response.json()
channels: dict[str, int] = {}
channels: dict[str, dict[str, Any]] = {}
for item in data:
raw = item.get("channel")
job_id = item.get("job_id")
@@ -46,7 +48,13 @@ class TelegramListener:
except ValueError:
logger.warning("Skip invalid channel in subscription: %r", raw)
continue
channels.setdefault(key, int(job_id))
channels.setdefault(
key,
{
"job_id": int(job_id),
"source_config": dict(item.get("source_config") or {}),
},
)
return channels
async def refresh_subscriptions(self) -> None:
@@ -73,10 +81,37 @@ class TelegramListener:
", ".join(sorted(channels)) or "(none)",
)
async def _ingest_post(self, channel: str, post) -> None:
def _build_event(self, channel: str, post) -> dict | None:
sub = self._channels.get(normalize_channel(channel)) or {}
cfg = sub.get("source_config") or {}
extract_mode = cfg.get("extract_mode") or "heuristic"
text = (post.text or "").strip()
if extract_mode == "profile" and cfg.get("heuristic_profile"):
return extract_with_profile(
text,
cfg["heuristic_profile"],
source_url=post.url,
source_type="telegram",
extra_metadata={
"channel": post.channel,
"message_id": post.id,
"post_date": post.date.isoformat() if post.date else None,
"listener": True,
},
)
# llm jobs fall back to heuristic in listener (LLM is batch-only by design)
record = parse_event_post(post)
event = event_record_to_ingest(record)
job_id = self._channels.get(normalize_channel(channel))
return event_record_to_ingest(record)
async def _ingest_post(self, channel: str, post) -> None:
event = self._build_event(channel, post)
if event is None:
logger.debug("Listener skip %s (profile required fields missing)", post.url)
return
sub = self._channels.get(normalize_channel(channel)) or {}
job_id = sub.get("job_id")
async with httpx.AsyncClient(timeout=60.0) as http:
response = await http.post(
@@ -1,20 +1,39 @@
"""Единое подключение Telethon для listener и batch-заданий."""
import logging
from pathlib import Path
from telethon import TelegramClient
from telethon.errors import AuthKeyUnregisteredError, SessionPasswordNeededError
from workers.sources.telegram_settings import create_client, get_api_credentials, get_session_path
from workers.sources.telegram_settings import (
create_client,
describe_connection,
get_api_credentials,
get_session_path,
)
from workers.sources.telegram_client import TelegramAuthError, TelegramConfigError
from workers.sources.vpn_config import proxy_fingerprint
logger = logging.getLogger("cp-worker.telegram-session")
_shared_client: TelegramClient | None = None
_proxy_fingerprint: str | None = None
async def get_shared_client() -> TelegramClient:
global _shared_client
global _shared_client, _proxy_fingerprint
fp = proxy_fingerprint()
if _shared_client is not None and _shared_client.is_connected():
return _shared_client
if fp == _proxy_fingerprint:
return _shared_client
logger.info(
"VPN proxy changed (%s → %s); reconnecting Telethon",
_proxy_fingerprint,
fp,
)
await close_shared_client()
try:
api_id, api_hash = get_api_credentials()
@@ -29,6 +48,7 @@ async def get_shared_client() -> TelegramClient:
)
client = create_client(session_path, api_id, api_hash)
logger.info("Connecting Telethon %s", describe_connection())
try:
await client.connect()
if not await client.is_user_authorized():
@@ -45,11 +65,13 @@ async def get_shared_client() -> TelegramClient:
) from exc
_shared_client = client
_proxy_fingerprint = fp
return client
async def close_shared_client() -> None:
global _shared_client
global _shared_client, _proxy_fingerprint
if _shared_client is not None:
await _shared_client.disconnect()
_shared_client = None
_proxy_fingerprint = None
@@ -28,7 +28,8 @@ def get_api_credentials() -> tuple[int, str]:
return api_id, api_hash
def get_proxy() -> tuple | None:
def get_env_proxy() -> tuple | None:
"""Legacy TELEGRAM_PROXY_* env fallback (used when CA VPN is off)."""
proxy_type = os.environ.get("TELEGRAM_PROXY_TYPE", "").strip().lower()
if not proxy_type or proxy_type == "none":
return None
@@ -49,6 +50,15 @@ def get_proxy() -> tuple | None:
return proxy_type, host, port
def get_proxy() -> tuple | None:
try:
from workers.sources.vpn_config import telethon_proxy_tuple
return telethon_proxy_tuple()
except Exception:
return get_env_proxy()
def describe_connection() -> str:
proxy = get_proxy()
if not proxy:
@@ -0,0 +1,124 @@
"""Fetch VPN settings from CA /internal/vpn (cached)."""
from __future__ import annotations
import logging
import os
import time
from typing import Any
import httpx
from contracts.vpn import VpnSettings
logger = logging.getLogger("cp-worker.vpn")
CA_API_URL = os.getenv("CA_API_URL", "http://ca-api:8000").rstrip("/")
INTERNAL_TOKEN = os.getenv("INTERNAL_TOKEN", "dev-internal-token")
CACHE_TTL = float(os.getenv("VPN_CONFIG_CACHE_SECONDS", "30"))
CP_VPN_HOST = os.getenv("CP_VPN_HOST", "cp-vpn")
CP_VPN_PORT = int(os.getenv("CP_VPN_PORT", "1080"))
_cache: VpnSettings | None = None
_cache_at: float = 0.0
_cache_failed: bool = False
def invalidate_vpn_cache() -> None:
global _cache, _cache_at, _cache_failed
_cache = None
_cache_at = 0.0
_cache_failed = False
def fetch_vpn_settings(*, force: bool = False) -> VpnSettings | None:
"""Return VPN settings from CA, or None if unreachable / unset."""
global _cache, _cache_at, _cache_failed
now = time.monotonic()
if (
not force
and _cache is not None
and (now - _cache_at) < CACHE_TTL
and not _cache_failed
):
return _cache
try:
with httpx.Client(timeout=10.0) as client:
response = client.get(
f"{CA_API_URL}/internal/vpn",
headers={"X-Internal-Token": INTERNAL_TOKEN},
)
response.raise_for_status()
raw: dict[str, Any] = response.json()
settings = VpnSettings.model_validate(raw)
_cache = settings
_cache_at = now
_cache_failed = False
return settings
except Exception:
logger.exception("Failed to load VPN settings from CA")
_cache_failed = True
_cache_at = now
# Keep stale cache if present
return _cache
def proxy_fingerprint(settings: VpnSettings | None = None) -> str:
cfg = settings if settings is not None else fetch_vpn_settings()
if cfg is None or not cfg.proxies_source("telegram"):
# Env fallback fingerprint
from workers.sources.telegram_settings import get_env_proxy
env = get_env_proxy()
return f"env:{env!r}"
if cfg.mode == "subscription":
return f"sub:{CP_VPN_HOST}:{CP_VPN_PORT}:{cfg.subscription_url}"
return (
f"{cfg.mode}:{cfg.host}:{cfg.port}:"
f"{cfg.username or ''}:{'*' if cfg.password else ''}"
)
def telethon_proxy_tuple(settings: VpnSettings | None = None) -> tuple | None:
"""Telethon-compatible proxy tuple for telegram traffic, or None."""
cfg = settings if settings is not None else fetch_vpn_settings()
if cfg is not None and cfg.proxies_source("telegram"):
if cfg.mode == "subscription":
return ("socks5", CP_VPN_HOST, CP_VPN_PORT)
if cfg.mode in ("socks5", "http") and cfg.host and cfg.port:
if cfg.username:
return (
cfg.mode,
cfg.host,
int(cfg.port),
True,
cfg.username,
cfg.password or "",
)
return (cfg.mode, cfg.host, int(cfg.port))
return None
from workers.sources.telegram_settings import get_env_proxy
return get_env_proxy()
def http_proxy_url(source_type: str, settings: VpnSettings | None = None) -> str | None:
"""HTTP(S)/SOCKS proxy URL for httpx / browsers, or None."""
cfg = settings if settings is not None else fetch_vpn_settings()
if cfg is None or not cfg.proxies_source(source_type):
return None
if cfg.mode == "subscription":
return f"socks5://{CP_VPN_HOST}:{CP_VPN_PORT}"
if cfg.mode in ("socks5", "http") and cfg.host and cfg.port:
auth = ""
if cfg.username:
from urllib.parse import quote
user = quote(cfg.username, safe="")
password = quote(cfg.password or "", safe="")
auth = f"{user}:{password}@"
return f"{cfg.mode}://{auth}{cfg.host}:{int(cfg.port)}"
return None
+218
View File
@@ -0,0 +1,218 @@
"""Declarative heuristic profile: schema + static rule interpreter (CA preview + CP runtime)."""
from __future__ import annotations
import re
from typing import Any, Literal
from pydantic import BaseModel, Field, field_validator, model_validator
TARGET_FIELDS: tuple[str, ...] = (
"title",
"description",
"locality",
"event_date",
"coords",
"topic",
"region",
)
COORDS_RE = re.compile(r"(-?\d{1,3}\.\d+)\s*,\s*(-?\d{1,3}\.\d+)")
Strategy = Literal["regex", "line", "after_marker", "between", "full_text", "literal"]
class FieldRule(BaseModel):
strategy: Strategy
# regex: named or group(1); line: 0-based index; after_marker/between: markers
pattern: str | None = None
group: int = 1
line_index: int | None = None
marker: str | None = None
end_marker: str | None = None
value: str | None = None # literal
flags: str = "" # e.g. "im" → re.I|re.M
strip: bool = True
@field_validator("pattern", "marker", "end_marker", "value", mode="before")
@classmethod
def empty_to_none(cls, value: Any) -> Any:
if value is None:
return None
if isinstance(value, str) and not value.strip():
return None
return value
class HeuristicProfile(BaseModel):
version: Literal[1] = 1
fields: dict[str, FieldRule] = Field(default_factory=dict)
required_fields: list[str] = Field(default_factory=list)
notes: str = ""
@field_validator("required_fields")
@classmethod
def known_required(cls, value: list[str]) -> list[str]:
seen: list[str] = []
unknown: list[str] = []
for name in value:
if name not in TARGET_FIELDS:
unknown.append(name)
continue
if name not in seen:
seen.append(name)
if unknown:
raise ValueError(f"Unknown required_fields: {sorted(set(unknown))}")
return seen
@model_validator(mode="after")
def known_fields_only(self) -> "HeuristicProfile":
unknown = set(self.fields) - set(TARGET_FIELDS)
if unknown:
raise ValueError(f"Unknown profile fields: {sorted(unknown)}")
return self
def target_field_specs() -> list[dict[str, str]]:
"""Fixed target table for profile UI (roadmap: custom tables later)."""
return [
{"name": "title", "type": "string", "description": "Short event title"},
{"name": "description", "type": "string", "description": "Event summary / body"},
{"name": "locality", "type": "string", "description": "Place / settlement name"},
{
"name": "event_date",
"type": "string",
"description": "Date as DD.MM.YYYY or YYYY-MM-DD",
},
{
"name": "coords",
"type": "string",
"description": "Latitude, longitude if present",
},
{"name": "topic", "type": "string", "description": "Short topic tag"},
{"name": "region", "type": "string", "description": "Region (optional)"},
]
def _compile_flags(flags: str) -> int:
mapping = {
"i": re.IGNORECASE,
"m": re.MULTILINE,
"s": re.DOTALL,
}
result = 0
for ch in (flags or "").lower():
result |= mapping.get(ch, 0)
return result
def _apply_rule(text: str, rule: FieldRule) -> str:
raw = text or ""
value = ""
if rule.strategy == "literal":
value = rule.value or ""
elif rule.strategy == "full_text":
value = raw
elif rule.strategy == "line":
lines = raw.splitlines()
idx = 0 if rule.line_index is None else rule.line_index
if 0 <= idx < len(lines):
value = lines[idx]
elif rule.strategy == "regex":
if not rule.pattern:
return ""
match = re.search(rule.pattern, raw, _compile_flags(rule.flags))
if match:
try:
value = match.group(rule.group)
except IndexError:
value = match.group(0)
elif rule.strategy == "after_marker":
marker = rule.marker or ""
if not marker:
return ""
pos = raw.find(marker)
if pos < 0:
return ""
start = pos + len(marker)
rest = raw[start:]
if rule.end_marker:
end = rest.find(rule.end_marker)
value = rest[:end] if end >= 0 else rest
elif rule.pattern:
match = re.search(rule.pattern, rest, _compile_flags(rule.flags))
if match:
try:
value = match.group(rule.group)
except IndexError:
value = match.group(0)
else:
# first non-empty line after marker
for line in rest.splitlines():
if line.strip():
value = line
break
elif rule.strategy == "between":
start_m = rule.marker or ""
end_m = rule.end_marker or ""
if not start_m or not end_m:
return ""
start = raw.find(start_m)
if start < 0:
return ""
start += len(start_m)
end = raw.find(end_m, start)
if end < 0:
return ""
value = raw[start:end]
if rule.strip:
value = value.strip()
return value
def apply_profile(text: str, profile: HeuristicProfile | dict[str, Any]) -> dict[str, str]:
"""Apply static rules to post text → string field map (no LLM)."""
if isinstance(profile, dict):
profile = HeuristicProfile.model_validate(profile)
result: dict[str, str] = {name: "" for name in TARGET_FIELDS}
for name, rule in profile.fields.items():
result[name] = _apply_rule(text, rule)
return result
def parse_coords(raw: str) -> tuple[float | None, float | None]:
match = COORDS_RE.search(raw or "")
if not match:
return None, None
return float(match.group(1)), float(match.group(2))
def _field_is_filled(name: str, value: str) -> bool:
if not (value or "").strip():
return False
if name == "coords":
lat, lng = parse_coords(value)
return lat is not None and lng is not None
return True
def match_profile(
text: str,
profile: HeuristicProfile | dict[str, Any],
) -> tuple[bool, dict[str, str], list[str]]:
"""Apply profile and report whether required_fields are filled.
Empty required_fields means no gate (legacy profiles ingest every post).
"""
if isinstance(profile, dict):
profile = HeuristicProfile.model_validate(profile)
fields = apply_profile(text, profile)
missing = [
name
for name in profile.required_fields
if not _field_is_filled(name, fields.get(name, ""))
]
return (not missing, fields, missing)
+86
View File
@@ -0,0 +1,86 @@
"""Reusable LLM extract profile (CA storage + flattened into TelegramSourceConfig)."""
from __future__ import annotations
from typing import Any
from pydantic import BaseModel, Field, field_validator, model_validator
from contracts.heuristic_profile import TARGET_FIELDS
# LLM defaults omit region (same as workers/llm_extract.DEFAULT_EXTRACT_SCHEMA).
LLM_SCHEMA_FIELDS: tuple[str, ...] = (
"title",
"locality",
"event_date",
"description",
"coords",
"topic",
)
DEFAULT_EXTRACT_SCHEMA: dict[str, str] = {
"title": "string — short event title",
"locality": "string — place / settlement name",
"event_date": "string — date as DD.MM.YYYY or YYYY-MM-DD if known",
"description": "string — concise event summary",
"coords": "string — latitude, longitude if present else empty",
"topic": "string — short topic tag",
}
DEFAULT_INSTRUCTION = (
"Extract structured military/news event fields from the text. "
"If the text is not an event, return is_event=false. "
"Respond with a single JSON object only."
)
class LlmProfile(BaseModel):
instruction: str | None = None
extract_schema: dict[str, str] = Field(default_factory=lambda: dict(DEFAULT_EXTRACT_SCHEMA))
required_fields: list[str] = Field(default_factory=list)
# When true, LLM may return multiple events per post (array "events")
multi_event: bool = False
@field_validator("instruction", mode="before")
@classmethod
def empty_instruction_to_none(cls, value: Any) -> Any:
if value is None:
return None
if isinstance(value, str) and not value.strip():
return None
return value
@field_validator("required_fields")
@classmethod
def known_required(cls, value: list[str]) -> list[str]:
seen: list[str] = []
unknown: list[str] = []
for name in value:
if name not in TARGET_FIELDS:
unknown.append(name)
continue
if name not in seen:
seen.append(name)
if unknown:
raise ValueError(f"Unknown required_fields: {sorted(set(unknown))}")
return seen
@model_validator(mode="after")
def known_schema_keys(self) -> "LlmProfile":
if not self.extract_schema:
raise ValueError("extract_schema must not be empty")
unknown = set(self.extract_schema) - set(TARGET_FIELDS)
if unknown:
raise ValueError(f"Unknown extract_schema fields: {sorted(unknown)}")
return self
def match_llm_required(fields: dict[str, Any], required_fields: list[str]) -> bool:
"""True if all required_fields are non-empty strings (empty required = no gate)."""
if not required_fields:
return True
for name in required_fields:
val = fields.get(name)
if val is None or not str(val).strip():
return False
return True
+41 -2
View File
@@ -10,16 +10,55 @@ from pydantic import BaseModel, Field, field_validator, model_validator
class TelegramSourceConfig(BaseModel):
channel: str = Field(min_length=1)
limit: int = Field(default=100, ge=1, le=1000)
# heuristic = legacy telegram_events parser; llm = DeepSeek structured extract
extract_mode: Literal["heuristic", "llm"] = "heuristic"
# heuristic = legacy telegram_events; llm = DeepSeek per post; profile = static rules
extract_mode: Literal["heuristic", "llm", "profile"] = "heuristic"
extract_schema: dict[str, str] | None = None
instruction: str | None = None
heuristic_profile: dict | None = None
# Post-extract gate for extract_mode=llm (empty = only is_event filter)
required_fields: list[str] = Field(default_factory=list)
# Split one post into N events when extract_mode=llm (opt-in)
multi_event: bool = False
sample_post: str | None = None # audit / re-generate; not required at runtime
@field_validator("channel")
@classmethod
def strip_channel(cls, value: str) -> str:
return value.strip()
@field_validator("required_fields")
@classmethod
def known_required(cls, value: list[str]) -> list[str]:
from contracts.heuristic_profile import TARGET_FIELDS
seen: list[str] = []
unknown: list[str] = []
for name in value:
if name not in TARGET_FIELDS:
unknown.append(name)
continue
if name not in seen:
seen.append(name)
if unknown:
raise ValueError(f"Unknown required_fields: {sorted(set(unknown))}")
return seen
@model_validator(mode="after")
def validate_extract_mode(self) -> "TelegramSourceConfig":
if self.extract_mode == "profile":
if not self.heuristic_profile:
raise ValueError("heuristic_profile required when extract_mode=profile")
from contracts.heuristic_profile import HeuristicProfile
HeuristicProfile.model_validate(self.heuristic_profile)
elif self.extract_mode == "llm" and self.extract_schema is not None:
from contracts.heuristic_profile import TARGET_FIELDS
unknown = set(self.extract_schema) - set(TARGET_FIELDS)
if unknown:
raise ValueError(f"Unknown extract_schema fields: {sorted(unknown)}")
return self
class Crawl4AISourceConfig(BaseModel):
urls: list[str] = Field(min_length=1)
+71
View File
@@ -0,0 +1,71 @@
"""VPN / proxy settings shared by CA admin UI and CP workers."""
from __future__ import annotations
from typing import Literal
from pydantic import BaseModel, Field, field_validator, model_validator
from contracts.queues import known_source_types
VpnMode = Literal["subscription", "socks5", "http"]
VPN_MODES: tuple[str, ...] = ("subscription", "socks5", "http")
class VpnSettings(BaseModel):
enabled: bool = False
mode: VpnMode = "subscription"
subscription_url: str | None = None
subscription_interval_seconds: int = Field(default=3600, ge=60, le=86400)
host: str | None = None
port: int | None = Field(default=None, ge=1, le=65535)
username: str | None = None
password: str | None = None
proxied_source_types: list[str] = Field(default_factory=list)
@field_validator("subscription_url", "host", "username", "password", mode="before")
@classmethod
def empty_to_none(cls, value: object) -> object:
if value is None:
return None
if isinstance(value, str):
stripped = value.strip()
return stripped or None
return value
@field_validator("proxied_source_types")
@classmethod
def validate_source_types(cls, value: list[str]) -> list[str]:
known = set(known_source_types())
cleaned: list[str] = []
seen: set[str] = set()
for raw in value or []:
item = str(raw).strip()
if not item or item in seen:
continue
if item not in known:
raise ValueError(
f"Unknown source_type {item!r}. "
f"Allowed: {', '.join(sorted(known))}"
)
seen.add(item)
cleaned.append(item)
return cleaned
@model_validator(mode="after")
def require_fields_when_enabled(self) -> "VpnSettings":
if not self.enabled:
return self
if self.mode == "subscription":
if not self.subscription_url:
raise ValueError("subscription_url required when mode=subscription")
elif self.mode in ("socks5", "http"):
if not self.host:
raise ValueError(f"host required when mode={self.mode}")
if self.port is None:
raise ValueError(f"port required when mode={self.mode}")
return self
def proxies_source(self, source_type: str) -> bool:
return self.enabled and source_type in self.proxied_source_types
+34 -3
View File
@@ -22,11 +22,16 @@ services:
build:
context: .
dockerfile: centers/analytics/api/Dockerfile
env_file:
- .env
environment:
DATABASE_URL: postgresql://mapmil:mapmil@ca-db:5432/mapmil
REDIS_URL: redis://redis:6379/0
INTERNAL_TOKEN: dev-internal-token
TEST_PI_API_KEY: test-pi-api-key-change-me
ADMIN_USER: ${ADMIN_USER:-admin}
ADMIN_PASSWORD: ${ADMIN_PASSWORD:-change-me}
ADMIN_JWT_SECRET: ${ADMIN_JWT_SECRET:-change-me-jwt-secret}
volumes:
- ca_uploads:/data
depends_on:
@@ -44,6 +49,17 @@ services:
- ca-api
restart: unless-stopped
cp-vpn:
build: ./centers/parsing/vpn
environment:
CA_API_URL: http://ca-api:8000
INTERNAL_TOKEN: ${INTERNAL_TOKEN:-dev-internal-token}
VPN_POLL_SECONDS: "15"
VPN_SOCKS_PORT: "1080"
depends_on:
- ca-api
restart: unless-stopped
cp-workers:
build:
context: .
@@ -57,16 +73,25 @@ services:
TELEGRAM_SESSION_PATH: /data/telegram.session
TELEGRAM_LISTENER_ENABLED: "true"
TELEGRAM_LISTENER_REFRESH_SECONDS: "60"
# Env fallback when VPN tab is off (host Happ/xray; socat 11080→10808).
# Overrides .env 127.0.0.1 which is wrong inside the container.
TELEGRAM_PROXY_HOST: host.docker.internal
TELEGRAM_PROXY_PORT: "11080"
CP_VPN_HOST: cp-vpn
CP_VPN_PORT: "1080"
ENABLED_ADAPTERS: telegram
WORKER_FAMILIES: telegram
CA_API_URL: http://ca-api:8000
REDIS_URL: redis://redis:6379/0
INTERNAL_TOKEN: dev-internal-token
INTERNAL_TOKEN: ${INTERNAL_TOKEN:-dev-internal-token}
extra_hosts:
- "host.docker.internal:host-gateway"
volumes:
- ./data:/data
depends_on:
- redis
- ca-api
- cp-vpn
restart: unless-stopped
cp-workers-web:
@@ -82,12 +107,15 @@ services:
TELEGRAM_LISTENER_ENABLED: "false"
ENABLED_ADAPTERS: crawl4ai
WORKER_FAMILIES: web
CP_VPN_HOST: cp-vpn
CP_VPN_PORT: "1080"
CA_API_URL: http://ca-api:8000
REDIS_URL: redis://redis:6379/0
INTERNAL_TOKEN: dev-internal-token
INTERNAL_TOKEN: ${INTERNAL_TOKEN:-dev-internal-token}
depends_on:
- redis
- ca-api
- cp-vpn
restart: unless-stopped
cp-workers-nlp:
@@ -101,12 +129,15 @@ services:
TELEGRAM_LISTENER_ENABLED: "false"
ENABLED_ADAPTERS: viina
WORKER_FAMILIES: nlp
CP_VPN_HOST: cp-vpn
CP_VPN_PORT: "1080"
CA_API_URL: http://ca-api:8000
REDIS_URL: redis://redis:6379/0
INTERNAL_TOKEN: dev-internal-token
INTERNAL_TOKEN: ${INTERNAL_TOKEN:-dev-internal-token}
depends_on:
- redis
- ca-api
- cp-vpn
restart: unless-stopped
volumes:
+35
View File
@@ -0,0 +1,35 @@
# Документация MapMil для разработчиков
Карта документов. Начните с обзора, затем углубитесь в нужный центр.
## С чего начать
1. [Обзор архитектуры](architecture-overview.md) — общая картина, затем детали
2. [Поток данных](data-flow.md) — от парсера до карты и ПИ
3. [Локальный запуск](local-dev.md) — `.env`, сессия, Docker
## По зонам
| Документ | Содержание |
|----------|------------|
| [architecture-overview.md](architecture-overview.md) | Платформа целиком: ЦА, ЦП, ПИ, границы, сервисы |
| [data-flow.md](data-flow.md) | E2E: admin → Redis → адаптер → ingest → карта → ПИ |
| [contracts.md](contracts.md) | Shared-схемы `contracts/` |
| [local-dev.md](local-dev.md) | Разработка и отладка локально |
| [ЦА: ARCHITECTURE](../centers/analytics/ARCHITECTURE.md) | API, модели, UI, scheduler |
| [ЦА: DISTRIBUTION (ПИ)](../centers/analytics/DISTRIBUTION.md) | External API для потребителей |
| [ЦП: ARCHITECTURE](../centers/parsing/ARCHITECTURE.md) | Адаптеры, очереди, Telegram listener |
## Корень репозитория
| Файл | Для кого |
|------|----------|
| [`README.md`](../README.md) | Быстрый старт и краткий API |
| [`AGENTS.md`](../AGENTS.md) | Ориентация для AI-агентов |
| [`.cursor/skills/add-parser-adapter/`](../.cursor/skills/add-parser-adapter/) | Чеклист нового `source_type` |
## Принцип
- **Contracts first** — изменение формы данных или маршрутизации начинается в `contracts/`.
- **Одна зона за раз** — не смешивать контракты + UI + инфра в одном несвязанном изменении.
- **ЦП не трогает БД ЦА** — только Redis и `POST/PATCH /internal/*`.
+297
View File
@@ -0,0 +1,297 @@
# Обзор архитектуры MapMil
Документ для разработчиков: сначала общая картина, затем устройство платформы подробнее.
Связанные материалы: [поток данных](data-flow.md), [контракты](contracts.md), [ЦА](../centers/analytics/ARCHITECTURE.md), [ЦП](../centers/parsing/ARCHITECTURE.md), [ПИ](../centers/analytics/DISTRIBUTION.md).
---
## 1. Общими чертами
**MapMil** — monorepo платформы сбора и отображения событий:
```text
Источники → ЦП (парсинг) → ЦА (хранение + карта + admin) → ПИ (внешние клиенты)
```
| Роль | Что это | Где код |
|------|---------|---------|
| **ЦА** (аналитика) | Единственный UI, PostgreSQL, ingest, admin, карта, distribution API | `centers/analytics/` |
| **ЦП** (парсинг) | Адаптеры источников без своей БД и без UI | `centers/parsing/workers/` |
| **ПИ** (потребители) | Не отдельный сервис: `GET /api/v1/events` в ЦА + ключи consumers | `centers/analytics/api/.../v1.py` |
| **Contracts** | Общие Pydantic-схемы очередей и payload | `contracts/` |
### Главный поток
```mermaid
flowchart LR
Admin[Admin UI / API]
Redis[(Redis queues)]
CP[CP adapters]
Src[Telegram / Web / NLP]
CA[(PostgreSQL)]
Map[Map UI]
PI[External API]
Admin -->|enqueue job| Redis
Redis -->|BLPOP| CP
CP --> Src
CP -->|POST /internal/ingest| CA
CA --> Map
CA --> PI
```
### Сценарий: создание парсера и взаимодействие
Объекты: **ЦА** (UI + API), **ЦП**, **ПИ**, **PostgreSQL**, **Redis**. Кто кому что передаёт при создании парсера и дальнейшем чтении данных.
```mermaid
sequenceDiagram
participant UI as ЦА UI
participant CA as ЦА API
participant PG as PostgreSQL
participant R as Redis
participant CP as ЦП
participant PI as ПИ клиент
Note over UI,CP: Создание парсера и batch-прогон
UI->>CA: POST /admin/jobs<br/>source_type + source_config
CA->>PG: INSERT ParseJob status=queued
CA->>R: RPUSH cp:jobs:family<br/>JobPayload job_id type config
R->>CP: BLPOP задание
CP->>CA: PATCH /internal/jobs/id<br/>status=running
CA->>PG: UPDATE ParseJob status
CP->>CP: fetch источника<br/>адаптер → события
CP->>CA: POST /internal/ingest<br/>events IngestEventItem
CA->>PG: INSERT Event<br/>дедуп по source_url
CA->>PG: sync MapObject<br/>если есть coords
CP->>CA: PATCH job completed/failed
CA->>PG: UPDATE ParseJob
Note over UI,PI: Чтение результата
UI->>CA: GET /api/map/objects
CA->>PG: SELECT map/events
CA-->>UI: точки на карте
PI->>CA: GET /api/v1/events<br/>Bearer API key
CA->>PG: SELECT events<br/>фильтры Consumer
CA-->>PI: срез событий
```
| От | Кому | Что передаётся |
|----|------|----------------|
| ЦА UI → ЦА API | Конфиг парсера (`source_type`, `source_config`) |
| ЦА API → PostgreSQL | `ParseJob`, затем `Event` / `MapObject` |
| ЦА API → Redis | `JobPayload` (`job_id`, type, config) |
| Redis → ЦП | То же задание (`BLPOP`) |
| ЦП → ЦА API | Статус job + пакет событий (ingest) |
| ЦА API → ЦА UI | Объекты карты |
| ЦА API → ПИ | Отфильтрованный срез событий |
### Жёсткие границы
1. ЦП **никогда** не подключается к PostgreSQL ЦА.
2. Связь ЦА ↔ ЦП: Redis (`cp:jobs:{family}`) + HTTP `/internal/*` с `X-Internal-Token`.
3. Весь фронтенд — только ЦА (`ca-frontend` на порту **8080**).
4. Дедуп событий в ЦА по стабильному `source_url`.
5. Секреты (`.env`, `data/telegram.session`) не коммитятся.
### Docker-сервисы одной командой
| Сервис | Назначение |
|--------|------------|
| `ca-db` | PostgreSQL |
| `redis` | Очереди заданий |
| `ca-api` | FastAPI ЦА |
| `ca-frontend` | Vue admin + карта |
| `cp-workers` | Telegram (batch + listener) |
| `cp-workers-web` | Crawl4AI |
| `cp-workers-nlp` | VIINA |
Запуск: `docker compose up --build` → http://localhost:8080
---
## 2. Подробнее
### 2.1. Структура monorepo
```text
MapMil/
├── centers/
│ ├── analytics/ # ЦА
│ │ ├── api/ # FastAPI
│ │ ├── frontend/ # Vue 3 + Leaflet
│ │ ├── ARCHITECTURE.md
│ │ └── DISTRIBUTION.md
│ └── parsing/ # ЦП
│ ├── ARCHITECTURE.md
│ └── workers/ # adapters + Telethon
├── contracts/ # ingest, jobs, sources, queues
├── docs/ # эта документация
├── data/ # telegram.session (локально)
├── docker-compose.yml
├── .env.example
├── README.md
└── AGENTS.md
```
### 2.2. Центр аналитики (ЦА)
**Ответственность:** владение данными, UI, постановка jobs, приём ingest, отдача срезов ПИ.
**Backend** (`centers/analytics/api/`):
| Слой | Содержание |
|------|------------|
| Routers | `/api/*` (карта, объекты), `/admin/*`, `/internal/*`, `/api/v1/*` |
| Models | `Event`, `ParseJob`, `MapObject`, `ObjectMedia`, `Consumer`, `ConsumerFilter` |
| Services | enqueue → Redis, scheduler, ingest + map sync, analytics queries |
| Auth internal | `X-Internal-Token` |
| Auth ПИ | Bearer API key (SHA-256 hash в БД) |
**Frontend** (`centers/analytics/frontend/`):
| Маршрут | Экран |
|---------|-------|
| `/` | Карта событий и ручных объектов |
| `/parsers` | CRUD парсеров (`telegram` / `crawl4ai` / `viina`) |
| `/events` | Список событий |
| `/analytics` | KPI и таймлайны |
| `/consumers` | Подписчики ПИ |
Nginx проксирует `/api/`, `/admin/`, `/internal/` на `ca-api:8000`.
Подробности: [centers/analytics/ARCHITECTURE.md](../centers/analytics/ARCHITECTURE.md).
### 2.3. Центр парсинга (ЦП)
**Ответственность:** забрать данные из внешнего источника и вернуть список `IngestEventItem`.
Адаптерный контракт:
```python
async def run(job_id, source_config, *, ctx) -> tuple[list[dict], str | None]
```
| source_type | Очередь | Контейнер |
|-------------|---------|-----------|
| `telegram` | `cp:jobs:telegram` | `cp-workers` |
| `crawl4ai` | `cp:jobs:web` | `cp-workers-web` |
| `viina` | `cp:jobs:nlp` | `cp-workers-nlp` |
Дополнительно для Telegram: **real-time listener** (Telethon `NewMessage` / `Album`) параллельно с batch, если `TELEGRAM_LISTENER_ENABLED=true`.
Опционально: `extract_mode: llm` (DeepSeek) для telegram/crawl4ai в batch.
**Профиль + канал:** UI `/parser-profiles` (CRUD, Generate/Preview для heuristic; instruction/schema/Preview для llm) и `/channels`; связка на `/parsers` → `ParseJob` с FK. При enqueue ЦА flatten по `ParserProfile.kind`:
- `heuristic` → `extract_mode=profile` + `heuristic_profile` (runtime без LLM; listener поддерживает);
- `llm` → `extract_mode=llm` + `extract_schema` / `instruction` / `required_fields` (DeepSeek на каждый пост в batch; listener пока fallback на heuristic).
Целевые поля — фиксированная схема Event/`IngestEventItem`. Кастомные пользовательские таблицы — roadmap.
Подробности: [centers/parsing/ARCHITECTURE.md](../centers/parsing/ARCHITECTURE.md), [centers/analytics/ARCHITECTURE.md](../centers/analytics/ARCHITECTURE.md).
### 2.4. Distribution (ПИ)
Отдельного центра нет. Внешний клиент:
1. Получает API-ключ в UI «ПИ» / admin API.
2. Вызывает `GET /api/v1/events` с `Authorization: Bearer <key>`.
3. Получает события с учётом фильтров consumer (регионы, темы, `date_from`).
Подробности: [centers/analytics/DISTRIBUTION.md](../centers/analytics/DISTRIBUTION.md).
### 2.5. Contracts
Единый источник правды для формы обмена:
| Модуль | Назначение |
|--------|------------|
| `contracts/jobs.py` | `JobPayload` в Redis |
| `contracts/ingest.py` | `IngestEventItem` / payload ingest |
| `contracts/sources.py` | Валидация `source_config` по `source_type` |
| `contracts/queues.py` | `SOURCE_FAMILY` → ключ очереди |
| `contracts/heuristic_profile.py` | Статичный профиль конструктора + `apply_profile` |
| `contracts/llm_profile.py` | Reusable LLM instruction/schema + `match_llm_required` |
Правило: меняете форму события / конфиг источника / очередь — сначала `contracts/`, потом CA/CP/UI.
Подробности: [contracts.md](contracts.md).
### 2.6. Данные в PostgreSQL (ядро)
```mermaid
erDiagram
ParseJob ||--o{ Event : produces
Event ||--o| MapObject : may_create
MapObject ||--o{ ObjectMedia : has
Consumer ||--o| ConsumerFilter : has
ParseJob {
int id
string source_type
json source_config
int interval_seconds
bool is_active
string status
}
Event {
int id
string source_url UK
string source_type
float latitude
float longitude
string region
string topic
}
MapObject {
int id
int event_id FK
float latitude
float longitude
}
Consumer {
int id
string name
string api_key_hash
}
```
- **ParseJob** — конфигурация парсера и статус batch-прогона.
- **Event** — нормализованное событие; уникальность по `source_url`.
- **MapObject** — точка на карте (из события с координатами или ручная).
- **Consumer** — внешний подписчик ПИ.
### 2.7. Планировщик
В `ca-api` фоновый scheduler (~каждые 30 с):
- берёт активные `ParseJob` со статусом `completed` / `failed`;
- если прошло ≥ `interval_seconds` с `last_run_at` — снова ставит в Redis (`queued`).
Так работают периодические парсеры без cron снаружи.
### 2.8. Секреты и тома
| Артефакт | Где | Заметка |
|----------|-----|---------|
| `.env` | корень | `TELEGRAM_*`, опционально `DEEPSEEK_*` |
| `data/telegram.session` | volume `./data` → `/data` в `cp-workers` | Telethon SQLite |
| `pgdata` | Docker volume | PostgreSQL |
| `ca_uploads` | Docker volume | медиа объектов карты |
---
## 3. Куда смотреть при типовых задачах
| Задача | Документ / код |
|--------|----------------|
| Новый источник парсинга | skill `add-parser-adapter` + `contracts/sources.py` + `queues.py` |
| Изменить поля события | `contracts/ingest.py` → ingest CA → map/ПИ |
| Карта / UI | `centers/analytics/frontend/src/` |
| Admin API | `centers/analytics/api/app/routers/admin.py` |
| Ingest / дедуп | `centers/analytics/api/app/services/ingest.py` |
| Telegram listener | `centers/parsing/.../telegram_listener.py` |
| Ключи для внешних систем | [DISTRIBUTION.md](../centers/analytics/DISTRIBUTION.md) |
+99
View File
@@ -0,0 +1,99 @@
# Contracts — общие схемы
Каталог [`contracts/`](../contracts/) — shared truth для формы данных и маршрутизации между ЦА и ЦП.
Правило платформы: **сначала contracts**, потом реализация в CA/CP/UI.
---
## Модули
| Файл | Назначение |
|------|------------|
| [`jobs.py`](../contracts/jobs.py) | Контракт задания. Payload задания в Redis |
| [`ingest.py`](../contracts/ingest.py) |Контракт результата. Событие и пакет ingest |
| [`sources.py`](../contracts/sources.py) | Контракт настроек парсера. Схемы `source_config` по `source_type` |
| [`queues.py`](../contracts/queues.py) | Контракт доставки задания нужному воркеру. `source_type` → family → Redis key |
| [`heuristic_profile.py`](../contracts/heuristic_profile.py) | Статичные правила extract_mode=profile |
| [`llm_profile.py`](../contracts/llm_profile.py) | Reusable LLM instruction/schema + required_fields + `multi_event` |
ЦА при admin CRUD валидирует конфиг через `parse_source_config`.
ЦП адаптеры должны отдавать dict, совместимые с `IngestEventItem`.
> На практике CA дублирует часть DTO в `app/schemas.py` для FastAPI; при изменении формы сверяйте оба места и UI.
---
## JobPayload (`jobs.py`)
```python
class JobPayload(BaseModel):
job_id: int
source_type: str
source_config: dict = {}
```
Уходит в Redis как JSON. Семейство очереди выбирается по `source_type`, не по содержимому config.
---
## Ingest (`ingest.py`)
Ключевые поля `IngestEventItem`:
| Поле | Обязательность | Заметка |
|------|----------------|---------|
| `source_url` | да | Стабильный URL; дедуп в ЦА |
| `source_type` | да (часто default) | Должен соответствовать адаптеру |
| `raw_text` / `title` / `description` | нет | Текст события |
| `latitude` / `longitude` | нет | Без них точка на карту не создаётся |
| `event_date`, `locality`, `region`, `topic` | нет | Фильтры карты и ПИ |
| `tags`, `metadata` | нет | Расширения |
Пакет: `{ job_id?, events: [...] }`.
Listener добавляет флаг `listener: true` на стороне CA API (см. internal schemas).
---
## Source config (`sources.py`)
| source_type | Модель | Главные поля |
|-------------|--------|--------------|
| `telegram` | `TelegramSourceConfig` | `channel`, `limit`, `extract_mode` (`heuristic`\|`llm`\|`profile`), `heuristic_profile`, `extract_schema`, `instruction`, `required_fields` (llm gate) |
| `crawl4ai` | `Crawl4AISourceConfig` | `urls`, `extract_mode`, `extract_schema`, `domain_profile` |
| `viina` | `ViinaSourceConfig` | `urls` / `texts`, `input_mode` |
Реестр: `CONFIG_MODELS` + `parse_source_config(source_type, raw)`.
Профили в БД ЦА (`ParserProfile`):
| kind | Хранилище | Flatten в Redis |
|------|-----------|-----------------|
| `heuristic` | `heuristic_profile` (`contracts/heuristic_profile.py`) | `extract_mode=profile` |
| `llm` | `llm_profile` (`contracts/llm_profile.py`: instruction, extract_schema, required_fields, multi_event) | `extract_mode=llm` (+ `#eN` URLs if multi) |
`required_fields` — обязательные поля; пустой список = без фильтра (для llm остаётся только `is_event`). Канал + профиль живут в ЦА; CP получает только плоский `source_config`.
---
## Очереди (`queues.py`)
```text
SOURCE_FAMILY = {
"telegram": "telegram", # cp:jobs:telegram
"crawl4ai": "web", # cp:jobs:web
"viina": "nlp", # cp:jobs:nlp
}
```
- `queue_key_for_source(source_type)` — куда enqueue из ЦА.
- Legacy-ключ `cp:jobs` ещё может дренироваться telegram-воркерами (совместимость).
При добавлении нового `source_type` обязательно:
1. схема в `sources.py`;
2. запись в `SOURCE_FAMILY`;
3. адаптер + registry + (при необходимости) новый Docker-сервис;
4. поля формы в `ParsersView.vue`.
Чеклист: [`.cursor/skills/add-parser-adapter/`](../.cursor/skills/add-parser-adapter/).
+134
View File
@@ -0,0 +1,134 @@
# Поток данных MapMil
Пошаговый путь события от настройки парсера до карты и внешнего API.
См. также: [обзор архитектуры](architecture-overview.md), [ЦП](../centers/parsing/ARCHITECTURE.md).
---
## Схема end-to-end
```mermaid
sequenceDiagram
participant UI as Admin UI
participant CA as ca-api
participant DB as PostgreSQL
participant R as Redis
participant CP as cp-workers
participant Src as Source
participant Map as Map UI
participant PI as External client
UI->>CA: POST /admin/jobs
CA->>DB: INSERT ParseJob queued
CA->>R: RPUSH cp:jobs:family
R->>CP: BLPOP
CP->>CA: PATCH /internal/jobs/id running
CP->>Src: fetch / listen
Src-->>CP: raw posts/pages
CP->>CA: POST /internal/ingest
CA->>DB: INSERT Event if new source_url
CA->>DB: sync MapObject if coords
CA->>CP: 200 ingested/skipped
CP->>CA: PATCH job completed or failed
Map->>CA: GET /api/map/objects
PI->>CA: GET /api/v1/events Bearer
```
---
## 1. Создание парсера (ЦА)
1. Пользователь в UI `/parsers` или `POST /admin/jobs` передаёт:
- `source_type` (`telegram` | `crawl4ai` | `viina`);
- `source_config` (валидируется `contracts/sources.py`);
- опционально `interval_seconds`, `is_active`.
2. ЦА пишет строку `ParseJob` со статусом `queued`.
3. ЦА делает `RPUSH` в Redis-ключ `cp:jobs:{family}` (`contracts/queues.py`).
Примеры `source_config`:
```json
{ "channel": "example", "limit": 50, "extract_mode": "heuristic" }
```
```json
{ "urls": ["https://news.example/"], "extract_mode": "llm" }
```
---
## 2. Batch-обработка (ЦП)
1. Воркер нужного семейства делает `BLPOP` своей очереди.
2. `PATCH /internal/jobs/{id}?status=running`.
3. Registry выбирает адаптер по `source_type` → `adapter.run(...)`.
4. Адаптер возвращает список dict в форме `IngestEventItem` (+ опциональный warning).
5. ЦП шлёт `POST /internal/ingest` с `{ job_id, events }`.
6. При успехе/ошибке — `PATCH` статуса `completed` / `failed`.
Если job попал не в ту очередь (чужой family), воркер **перекладывает** его в правильную очередь.
---
## 3. Real-time Telegram (параллельно)
Только `cp-workers` при `TELEGRAM_LISTENER_ENABLED=true`:
1. Listener опрашивает `GET /internal/listener/subscriptions` (активные telegram-jobs).
2. Telethon получает `NewMessage` / `Album`.
3. Пост парсится (эвристика) → один event → `POST /internal/ingest` с `listener: true`.
4. Статус `ParseJob` от listener **не** переводится в `completed` (в отличие от batch).
---
## 4. Ingest в ЦА
Сервис `centers/analytics/api/app/services/ingest.py`:
1. Для каждого item проверяет уникальность `source_url`.
2. Уже есть → skip (дедуп).
3. Нет → INSERT `Event`.
4. Если есть `latitude` / `longitude` → создать/обновить связанный `MapObject`.
5. Batch (`listener` не true) обновляет статус job и `last_run_at`.
---
## 5. Карта
1. UI периодически (около 30 с) и по фильтрам вызывает `GET /api/map/objects`.
2. Ответ — точки с данными события (даты, регион, тема, источник).
3. Ручные объекты карты живут в `map_objects` без обязательного `event_id` (CRUD через UI / `/api/objects`).
---
## 6. Внешний потребитель (ПИ)
1. В UI `/consumers` создаётся consumer + API-ключ (в БД хранится hash).
2. Клиент: `GET /api/v1/events` + `Authorization: Bearer <key>`.
3. ЦА применяет `ConsumerFilter` (регионы, темы, `date_from`) и отдаёт срез событий.
---
## 7. Периодический перезапуск
Scheduler в `ca-api` (~30 с):
- активный job;
- статус `completed` или `failed`;
- прошло ≥ `interval_seconds` с последнего запуска
→ снова `queued` + enqueue в Redis.
---
## Точки отказа (для отладки)
| Симптом | Куда смотреть |
|---------|----------------|
| Job висит в `queued` | Redis, нужный `cp-workers*`, `WORKER_FAMILIES` |
| Job `failed` | `last_error` в admin, логи контейнера адаптера |
| Событий 0 при успехе | Парсер/LLM не извлёк поля; дедуп по `source_url` |
| Нет точек на карте | Нет координат у Event; фильтры даты/региона в UI |
| 401 на `/api/v1/events` | Ключ, `is_active` consumer |
| Telegram auth error | `data/telegram.session`, `TELEGRAM_API_ID/HASH` |
+133
View File
@@ -0,0 +1,133 @@
# Локальная разработка
## Требования
- Docker + Docker Compose
- Файл `.env` (из `.env.example`)
- Для Telegram: `data/telegram.session` + `TELEGRAM_API_ID` / `TELEGRAM_API_HASH`
- Опционально: `DEEPSEEK_API_KEY` для `extract_mode: llm` (воркеры) и Generate профиля на `ca-api`
## Быстрый старт
```bash
cp .env.example .env
# заполнить TELEGRAM_* (и при необходимости DEEPSEEK_*)
mkdir -p data
# положить telegram.session в data/
docker compose up --build
```
UI: http://localhost:8080
Health: `curl http://localhost:8080/api/health`
## Сервисы
| Сервис | Порт снаружи | Когда нужен |
|--------|--------------|-------------|
| `ca-frontend` | 8080 | всегда (UI) |
| `ca-api` | внутренний 8000 | всегда |
| `ca-db` / `redis` | — | всегда |
| `cp-workers` | — | Telegram |
| `cp-workers-web` | — | Crawl4AI |
| `cp-workers-nlp` | — | VIINA |
| `cp-vpn` | внутренний 1080 | SOCKS из subscription (вкладка VPN) |
Пересборка одного воркера:
```bash
docker compose up --build cp-workers-web
```
Проверка compose-файла:
```bash
docker compose config
```
## Telegram-сессия
- Путь в контейнере: `/data/telegram.session` (`TELEGRAM_SESSION_PATH`)
- Локально: `./data` монтируется в `cp-workers`
- Файл **не** коммитить (`.gitignore`)
- API id/hash в `.env` должны совпадать с теми, под которыми создавалась сессия
## VPN / прокси (вкладка VPN)
Основной путь: UI → **VPN** (`/vpn`).
1. Включите «Прокси включён».
2. Способ **Subscription URL** — вставьте ссылку подписки (не коммитьте её в git).
3. Отметьте источники (`telegram`, при необходимости `crawl4ai` / `viina`).
4. Сохраните. Сервис `cp-vpn` (mihomo) подтянет подписку и отдаст SOCKS на `cp-vpn:1080`; воркеры читают `/internal/vpn`.
Режимы **SOCKS5** / **HTTP** задают host:port напрямую (без `cp-vpn` для Telegram).
### Fallback: локальный Happ/xray (без вкладки VPN)
Если VPN в UI выключен, воркеры используют `TELEGRAM_PROXY_*` из `.env`:
1. Остановите Telegram-воркер, чтобы не делить session-файл:
`docker compose stop cp-workers`
2. В `.env`: `TELEGRAM_PROXY_TYPE=socks5`, `TELEGRAM_PROXY_HOST=127.0.0.1`, `TELEGRAM_PROXY_PORT=10808` (Happ/xray).
3. Проброс SOCKS в Docker (xray слушает только `127.0.0.1:10808`):
```bash
socat TCP-LISTEN:11080,bind=0.0.0.0,fork,reuseaddr TCP:127.0.0.1:10808 &
```
Для auth из контейнера с host-сетью: `TELEGRAM_PROXY_HOST=127.0.0.1`.
4. Авторизация (интерактивно, код из Telegram):
```bash
docker run --rm -it --network host \
--env-file .env \
-e TELEGRAM_SESSION_PATH=/data/telegram.session \
-e TELEGRAM_PROXY_HOST=127.0.0.1 \
-v "$PWD/data:/data" \
-v "$PWD/scripts:/scripts:ro" \
mapmil-cp-workers \
python /scripts/telegram_auth.py
```
5. `docker compose up -d cp-workers`
Без сессии batch/listener Telegram не авторизуются; остальные адаптеры могут работать.
## Полезные curl
Создать telegram-парсер:
```bash
curl -X POST http://localhost:8080/admin/jobs \
-H 'Content-Type: application/json' \
-d '{"source_type":"telegram","source_config":{"channel":"example","limit":50}}'
```
Срез ПИ (подставьте ключ из seed / UI):
```bash
curl -H "Authorization: Bearer test-pi-api-key-change-me" \
"http://localhost:8080/api/v1/events"
```
## Логи
```bash
docker compose logs -f ca-api
docker compose logs -f cp-workers
docker compose logs -f cp-workers-web
```
## Типичные проблемы
| Проблема | Действие |
|----------|----------|
| Frontend 000 / нет контейнеров | `docker compose up -d` (без `--build`, если registry недоступен, но образы уже есть) |
| Job не берётся | Смотреть `ENABLED_ADAPTERS` / `WORKER_FAMILIES` нужного сервиса |
| LLM не работает | `DEEPSEEK_API_KEY` в `.env`, перезапуск `cp-workers` / `cp-workers-web` |
| Generate профиля 503 | `DEEPSEEK_API_KEY` в `.env` + `env_file` у `ca-api`, перезапуск `ca-api` |
| Изменения UI не видны | Пересобрать `ca-frontend` |
## Границы при разработке
- Не добавлять прямой доступ к БД из CP workers.
- Не коммитить `.env`, `data/`, API-ключи.
- Изменение формы данных — через `contracts/` (см. [contracts.md](contracts.md)).
+115
View File
@@ -0,0 +1,115 @@
#!/usr/bin/env python3
"""Интерактивная авторизация Telethon → data/telegram.session.
Запуск через образ cp-workers и host-сеть (локальный xray/Happ на 127.0.0.1:10808):
docker compose stop cp-workers
docker run --rm -it --network host \\
--env-file .env \\
-e TELEGRAM_SESSION_PATH=/data/telegram.session \\
-e TELEGRAM_PROXY_HOST=127.0.0.1 \\
-v \"$PWD/data:/data\" \\
-v \"$PWD/scripts:/scripts:ro\" \\
mapmil-cp-workers \\
python /scripts/telegram_auth.py
docker compose up -d cp-workers
"""
from __future__ import annotations
import asyncio
import os
import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[1]
WORKERS = ROOT / "centers" / "parsing" / "workers"
if WORKERS.exists() and str(WORKERS) not in sys.path:
sys.path.insert(0, str(WORKERS))
def _load_dotenv(path: Path) -> None:
if not path.is_file():
return
for line in path.read_text(encoding="utf-8").splitlines():
line = line.strip()
if not line or line.startswith("#") or "=" not in line:
continue
key, _, value = line.partition("=")
key = key.strip()
value = value.strip().strip("'").strip('"')
os.environ.setdefault(key, value)
async def main() -> int:
_load_dotenv(ROOT / ".env")
# Prefer repo-local session when running outside Docker /data mount.
if not os.environ.get("TELEGRAM_SESSION_PATH"):
local = ROOT / "data" / "telegram.session"
os.environ["TELEGRAM_SESSION_PATH"] = str(local)
from workers.sources.telegram_settings import (
create_client,
describe_connection,
get_api_credentials,
get_session_path,
)
api_id, api_hash = get_api_credentials()
session_path = get_session_path()
Path(session_path).parent.mkdir(parents=True, exist_ok=True)
print(f"Сессия: {session_path}")
print(f"Подключение: {describe_connection()}")
client = create_client(session_path, api_id, api_hash)
await client.connect()
if await client.is_user_authorized():
me = await client.get_me()
print(
f"Уже авторизовано: id={me.id} "
f"username={getattr(me, 'username', None) or '—'} "
f"phone={getattr(me, 'phone', None) or '—'}"
)
await client.disconnect()
return 0
print("Сессия не авторизована — вход в Telegram.")
phone = input("Номер телефона (+7...): ").strip()
if not phone:
print("Номер не указан.", file=sys.stderr)
await client.disconnect()
return 1
await client.send_code_request(phone)
code = input("Код из Telegram/SMS: ").strip()
try:
await client.sign_in(phone=phone, code=code)
except Exception as exc:
# 2FA
from telethon.errors import SessionPasswordNeededError
if not isinstance(exc, SessionPasswordNeededError):
raise
password = input("Пароль 2FA: ").strip()
await client.sign_in(password=password)
me = await client.get_me()
print(
f"Готово. Авторизован: id={me.id} "
f"username={getattr(me, 'username', None) or '—'} "
f"phone={getattr(me, 'phone', None) or '—'}"
)
print(f"Файл сессии: {session_path}")
await client.disconnect()
return 0
if __name__ == "__main__":
try:
raise SystemExit(asyncio.run(main()))
except KeyboardInterrupt:
print("\nОтменено.", file=sys.stderr)
raise SystemExit(130)