panoptes.data.search¶
panoptes.data.search ¶
MetadataUnavailableError ¶
Bases: RuntimeError
Some sequences' metadata could not be read, so the table is incomplete.
Carries what it knows rather than only that it failed: failures maps each
sequence id to the exception it raised, and partial is the table of the
sequences that did read. A caller who genuinely wants an incomplete result
takes it from here, which is the difference between choosing a partial
answer and being handed one.
Source code in src/panoptes/data/search.py
parse_duration ¶
A duration as a span and the direction it runs, if it says.
Accepts "90 days", "6 months before", "10 days before and after"
and a datetime.timedelta. The direction is None when the string does not
give one, which leaves the choice to duration_window; a negative
timedelta names one, since a window cannot run backward from its start.
Raises:
| Type | Description |
|---|---|
ValueError
|
if the string is not a duration this understands. The message lists the forms that work, because a duration that silently parsed as something else would move the search window without saying so. |
Source code in src/panoptes/data/search.py
duration_window ¶
duration_window(
duration: str | timedelta,
anchor: Timestamp,
default: str,
) -> tuple[pd.Timestamp, pd.Timestamp]
The (start, end) a duration marks out around anchor.
default is the direction to use when the duration does not name one. A
duration anchored on a start_date runs forward from it; one with no
start_date is anchored on now and runs backward, because a window in the
future holds no observations.
"10 days before and after" is why this returns a pair rather than an end
date: a window can sit on both sides of its anchor, and no single parsed
datetime can say so.
Source code in src/panoptes/data/search.py
as_utc ¶
wrap_degrees ¶
Signed separation in degrees, folded onto (-180, 180].
A mount at 359.9 degrees and a target at 0.1 are 0.2 degrees apart, not 359.8. Comparing raw degrees makes every cone spanning 0h RA silently return nothing.
Source code in src/panoptes/data/search.py
add_pointing ¶
Attach each sequence's mean pointing, and how far it wandered from it.
observations.parquet carries no coordinates: the index groups frames
into sequences, and RA-MNT/DEC-MNT are per-frame header readings, so
there is no single sequence coordinate to store. This derives one.
The mean is not enough on its own. Older observations drift
substantially over a night -- that is the point of measuring drift at all
(panoptes-pipeline improvement plan 3.2) -- so a sequence whose mean sits
just outside a search cone may still have spent half the night inside it.
Each sequence therefore also gets the largest deviation of any of its frames
from its own mean, and search_observations widens the cone by that amount
per sequence rather than by one fleet-wide fudge factor.
RA is averaged as an angle, via the mean unit vector, so a sequence straddling 0h averages to 0 rather than to 180.
Returns:
| Type | Description |
|---|---|
DataFrame
|
|
DataFrame
|
frame coordinates get nulls, and a null never matches a cone. |
Source code in src/panoptes/data/search.py
add_frame_facts ¶
Attach the per-frame header facts a sequence can be cut on.
observations.parquet groups frames into sequences and keeps what is
meaningful for a sequence, so ISO, airmass and the moon columns are not in
it -- each is a reading taken per frame. They are nonetheless four of the
nine cuts data contract 9 says benchmark selection has to make, and
answering "which sequences were shot at ISO 100 with the moon down" one
sequence at a time is the 12,439 document reads that issue is about. So they
are reduced here, in the same pass over frames.parquet that already
derives pointing.
A column no document supplied arrives as null rather than as a missing
column, so a cut on it selects nothing instead of raising. That matches
add_pointing: an index built over documents that never carried the field
is a valid index, and a query against it should come back empty rather than
refuse to run.
Returns:
| Type | Description |
|---|---|
DataFrame
|
|
Source code in src/panoptes/data/search.py
find_simultaneous ¶
find_simultaneous(
observations: DataFrame,
across: str = "camera_id",
min_overlap_minutes: float = 0.0,
same_field: bool = True,
) -> pd.DataFrame
Sequences of the same field shot at the same time by different bodies.
PAN007_d37295_20250407T061910 and PAN007_f6eb3d_20250407T061910 are
372 frames each, the same night, the same field, two different cameras on
one unit. A pair observed simultaneously is the control the photometry
rebuild compares against, because the sky was the same and the hardware was
not (panoptes/panoptes-data#14).
Overlap in time, not "the same night". The units are spread across
longitudes, so a calendar date is a different slice of an observing night
for each of them, and any UTC-based night key is wrong for somebody.
start_time and end_time are contract columns, so asking whether two
sequences were running at once needs no such convention and admits no
argument. Sequences whose extent the index cannot report are dropped, since
an unknown interval overlaps nothing that can be checked.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
observations
|
DataFrame
|
Sequences to pair, as |
required |
across
|
str
|
The column that must differ between the two sequences of a
pair -- |
'camera_id'
|
min_overlap_minutes
|
float
|
Discard pairs overlapping by less than this. The default keeps any overlap at all, including two sequences that merely touch. |
0.0
|
same_field
|
bool
|
Require both sequences to carry the same |
True
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
One row per pair, the columns of |
DataFrame
|
and |
DataFrame
|
the field the pair shares, and is null when the two sequences were on |
DataFrame
|
different fields -- |
DataFrame
|
both. A table with no pairs still has those columns. |
Raises:
| Type | Description |
|---|---|
ValueError
|
if |
Source code in src/panoptes/data/search.py
322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 | |
search_observations ¶
search_observations(
*,
by_name=None,
coords=None,
unit_id=None,
start_date=None,
end_date=None,
duration=None,
ra=None,
dec=None,
radius=10,
min_num_frames=1,
min_num_usable=None,
min_duration_minutes=None,
field_name=None,
camera_id=None,
query=None,
source=None,
ra_col="mount_ra",
dec_col="mount_dec",
) -> pd.DataFrame
Search PANOPTES observations.
The columns are the ones panoptes-pipeline declares -- num_frames,
num_usable, duration_minutes, sequence_sequence_id -- rather
than the Firestore summary's, which nothing produces.
See panoptes.data.documents.
A position is optional. With no coords, by_name, or ra/dec, the
search is all-sky and every other filter still applies
(panoptes/panoptes-data#14).
from astropy.coordinates import SkyCoord from panoptes.data.search import search_observations coords = SkyCoord.from_name('Andromeda Galaxy') start_date = '2019-01-01' end_date = '2019-12-31' search_results = search_observations(coords=coords, min_num_frames=10, ... start_date=start_date, end_date=end_date)
The result is a DataFrame you can further work with.¶
search_results.groupby(['unit_id', 'field_name']).num_usable.sum()
The benchmark criterion of data contract 9 -- 300 usable frames over three hours, anywhere in the sky -- is the whole surface at once:
benchmarks = search_observations(start_date='2024-01-01', ... min_num_usable=300, ... min_duration_minutes=180, ... query='moonfrac < 0.25 and airmass < 1.5')
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
by_name
|
str | None
|
If present, this will use the |
None
|
coords
|
`astropy.coordinates.SkyCoord`|None
|
A valid coordinate instance. |
None
|
ra
|
float | None
|
The RA position in degrees of the center of search.
Requires |
None
|
dec
|
float | None
|
The Dec position in degrees of the center of the search. |
None
|
radius
|
float
|
The search radius in degrees. Searches are done in a
square box, so this is half the length of the side of the box. Each
sequence's box is widened by that sequence's own pointing drift, so
an observation that wandered into the box is found; see
|
10
|
start_date
|
str|`datetime.datetime`|None
|
A valid datetime instance or |
None
|
end_date
|
str|`datetime.datetime`|None
|
A valid datetime instance or |
None
|
duration
|
str|`datetime.timedelta`|None
|
The length of the window,
instead of naming its far end: The window is anchored on
|
None
|
unit_id
|
str | list | None
|
A str or list of strs of unit_ids to include.
Default |
None
|
min_num_frames
|
int
|
Minimum number of frames the observation should
have, default 1. This counts frames the pipeline has a document for;
|
1
|
min_num_usable
|
int | None
|
Minimum number of frames the pipeline processed cleanly. This is the count a benchmark criterion means: one 372-frame sequence in the archive has 310 usable frames and 62 errors, and only this filter can tell them apart. |
None
|
min_duration_minutes
|
float | None
|
Minimum wall-clock span from the first frame to the last, which is what "3+ hours" asks for. A sequence whose span the index could not compute is excluded rather than assumed long enough. |
None
|
field_name
|
str | list | None
|
A str or list of strs of field names to
include. Default |
None
|
camera_id
|
str | list | None
|
A str or list of strs of camera uids to
include. Default |
None
|
query
|
str | None
|
A |
None
|
source
|
`pandas.DataFrame`|None
|
The table to search. If |
None
|
ra_col
|
str
|
The column to use for the RA, default 'mount_ra'. |
'mount_ra'
|
dec_col
|
str
|
The column to use for the Dec, default 'mount_dec'. |
'mount_dec'
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
if exactly one of |
Source code in src/panoptes/data/search.py
456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 | |
get_all_observations ¶
get_all_observations(
settings: SurveySettings | None = None,
index_root: Path | str | None = None,
) -> pd.DataFrame
Every sequence in the index, with its pointing attached.
Reads observations.parquet, the query surface panoptes-pipeline
builds by walking the documents it wrote.
Two things come for free with parquet. Dtypes are in the file, so
camera_serial_number comes back as the string 032071000633 rather
than being inferred into 3.207100e+10 (panoptes/panoptes-data#13). And
total_exptime is a sum over per-frame records, so it is populated for
the long sequences.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
SurveySettings | None
|
The settings to use. Defaults to reading the environment. |
None
|
index_root
|
Path | str | None
|
The directory holding the index files, overriding the settings for this call. |
None
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
pd.DataFrame: One row per sequence, with |
Raises:
| Type | Description |
|---|---|
DocumentsUnavailableError
|
if no index is configured or none is there. |
Source code in src/panoptes/data/search.py
get_metadata ¶
Read the per-frame documents of many sequences into one table.
Failing is the default, and a caller who wants only what could be read has to say so. Swallowing a per-sequence failure makes a run in which half the sequences failed indistinguishable from one that read the whole archive (panoptes/panoptes-data#13), and the rest of the package follows the same rule: partial data raises, it never silently shortens.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
observations
|
DataFrame
|
Rows carrying a sequence id, as |
required |
errors
|
str
|
|
'raise'
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
pd.DataFrame: The frames of every sequence, concatenated. Empty input
gives an empty table rather than a |
Raises:
| Type | Description |
|---|---|
MetadataUnavailableError
|
with |
ValueError
|
if |