[Save this search]

Status
All
   Fixed (7187)
  Closed (5379)
✓Open (2656)
   Won't Fix (545)
   Duplicate (297)
   Invalid (217)
   Not A Problem (195)
Issue type
✓All
  Issue (2284)
  PR (372)
Author relation
✓All
  None (2063)
  Member (669)
  Contributor (522)
  New contributor (74)
Created
✓All
  Past day (3)
  Past 2 days (9)
  Past 3 days (14)
  Past week (20)
  Past month (55)
  Past 3 months (116)
  Past 6 months (198)
  Past year (280)
Updated
✓All
  Past day (16)
  Past 2 days (28)
  Past 3 days (39)
  Past week (52)
  Past month (150)
  Past 3 months (221)
  Past 6 months (267)
  Past year (510)
Updated ago
✓All
  > 1 day ago (2640)
  > 2 days ago (2628)
  > 3 days ago (2617)
  > 1 week ago (2604)
  > 1 month ago (2506)
  > 3 months ago (2435)
  > 1 year ago (2146)
Comment count
✓All
  0 (580)
  1 (385)
  2 - 5 (824)
  6 - 10 (456)
  10 - 20 (312)
  > 20 (173)
Reaction count
✓All
  0 (2366)
  1 (174)
  2 - 5 (94)
  6 - 10 (16)
  10 - 20 (5)
  > 20 (1)
Review Requested
✓All
  jpountz (14)
  mikemccand (13)
  rmuir (6)
  iverase (4)
  romseygeek (4)
  benwtrent (4)
  dweiss (3)

See all 28...
Mentioned
✓All
  jpountz (97)
  mikemccand (86)
  rmuir (65)
  benwtrent (53)
  msokolov (45)
  uschindler (42)
  romseygeek (22)

See all 218...
Reviewed
✓All
  mikemccand (25)
  jpountz (23)
  rmuir (15)
  benwtrent (14)
  uschindler (10)
  dweiss (9)
  msokolov (9)

See all 64...
Commented
✓All
  asfimport (1366)
  github-actions[bot] (280)
  jpountz (136)
  mikemccand (126)
  rmuir (100)
  msokolov (95)
  benwtrent (82)

See all 316...
User
✓All
  asfimport (1805)
  github-actions[bot] (341)
  mikemccand (221)
  jpountz (204)
  msokolov (163)
  rmuir (143)
  benwtrent (120)

See all 483...
Last comment user
✓All
  asfimport (1316)
  github-actions[bot] (262)
  mikemccand (33)
  jpountz (31)
  rmuir (24)
  msokolov (23)
  benwtrent (17)

See all 175...
Draft
✓All
  No (307)
  Yes (65)
Component
✓All
  core (655)
  analysis (150)
  highlighter (46)
  facet (44)
  spatial (41)
  queryparser (29)
  sandbox (28)

See all 23...
Type
✓All
  enhancement (1208)
  bug (722)
  task (201)
  test (81)
  documentation (21)
Labels
✓All
  Stale (222)
  legacy-jira-fix-versio... (214)
  legacy-jira-fix-versio... (169)
  affects-version:4.0-ALPHA (81)
  tool:build (54)
  vector-based-search (48)
  affects-version:6.0 (37)

See all 156...
Commits?
✓All
  No (2656)
Reporter
✓All
  rmuir (269)
  mikemccand (156)
  jpountz (123)
  dsmiley (68)
  uschindler (52)
  iverase (44)
  romseygeek (41)

See all 769...
Assignee
✓All
  Unassigned (2415)
  mikemccand (35)
  uschindler (31)
  romseygeek (27)
  dsmiley (22)
  rmuir (15)
  jpountz (11)

See all 45...
  Filters: Status (Open),  Issue type,  Author relation,  Created,  Updated,  Updated ago,  Comment count,  Reaction count,  Review Requested,  Mentioned,  Reviewed,  Commented,  User,  Last comment user,  Draft,  Component,  Type,  Labels,  Commits?,  Reporter,  Assignee

#16794 PR: Speed up concurrent HNSW merges by joining the smaller graphs
1.9 hours ago  11 comments  0 votes  0 watches  LantaoJin, Pulkitg64, benwtrent, github-actions[bot], kaivalnp, mayya-sharipova, msokolov
Description ConcurrentHnswMerger never uses the join-set merge from 14331. It initializes the merged graph from the largest input graph and then inserts every other vector ... Only the serial MergingHnswGraphBuilder joins them.
    mayya-sharipova 1.9 hours ago:  nit: when the base graph has deletes, rebalanceGraph() runs on the calling thread while all workers ... The plans only read the graphs being joined, so they don't depend on the repair or the rebalance. They could share one invokeAll with the rebalance, e.g. invokeAll([rebalance, plan(g1), …, plan(gG) ...
    mayya-sharipova 1.9 hours ago:  Instead of this lock, could each worker get its own instance per joined graph, e.g. by having ... That would remove the lock. It costs a few small objects per worker and graph.

#16771 PR: Page-align vector data for all encodings
4 hours ago  7 comments  0 votes  0 watches  CH-Abhinav, Pulkitg64, github-actions[bot], goankur, kaivalnp
Description Fast-follow to 16705. Rescoring with full-precision vectors reads one whole vector per candidate. A 1024-dim float32 vector is exactly 4096 bytes, but Lucene99FlatVectorsWriter aligns the vector ...
    goankur 4 hours ago:  Folks, can a second pair of eyes take a look and merge it if there are no objections ? I'd really like this to be part of 11.0 release if its not too late.
    goankur 2.2 days ago:  Thanks @kaivalnp. Agreed, main does not rescore from bytes today, so for BYTE the padding only helps when HNSW reads ... Byte rescoring is planned.

#16805 PR: Collector manager join [DRAFT]
4.1 hours ago  8 comments  0 votes  0 watches  github-actions[bot], mkhludnev, msokolov
I took another stab at cutting over the usage in the join package from Collector to CollectorManager ... It seems to pass all the tests, but I don't have a high degree of confidence - the testing in this ... I wonder if anyone who uses this stuff could try this out and give some idea of performance impact? ...
    msokolov 4.1 hours ago:  This is not what I was assuming in the case of Avg -- it has *already* been divided. Instead we need the raw sums
    msokolov 4.3 hours ago:  I guess they get accumulated into joinValues (and scores) as a side effect of collection

#16804 PR: Cap the backoff between residency checks in MemorySegmentIndexInput prefetch
5.7 hours ago  0 comments  0 votes  0 watches  github-actions[bot], pauldoo
The MemorySegmentIndexInput prefetch logic contains a heuristic to issue madvise calls only when it ... We observed that after a long run of checks reporting data being loaded, the system fails to react ... This results in an extended period of low IO rate while the system is major faulting on every read. ...

#16145 PR: Delay and cap the prefetch backoff in MemorySegmentIndexInput
6 hours ago  33 comments  0 votes  0 watches  github-actions[bot], goankur, iprithv, jimczi, michaeljmarshall, mikemccand, navneet1v, neoremind, uschindler
Description The power-of-two back off in MemorySegmentIndexInput.prefetch() assumes consecutive ... This assumption works well for files that are already warm, but is too aggressive when files are ... The proposed solution makes two primary changes: 1.
    goankur 6 hours ago:  @michaeljmarshall @uschindler here is a full matrix from one box and one set of builds, so all ... Summary: 1. **Randomness is not needed.** Sampling by a hash of the page (BitMixer.mix64(offset >>> 12)) ...
    michaeljmarshall 1.1 days ago:  My initial goal with the bit mixing and counting local clone calls to prefetch was to spread out ... Then, @goankur suggested that wasn't ideal in the first case: > 1. shouldProbe(prefetchCount++) ... SegmentTermsEnum clones termsIn for each lookup, so every term lookup in every segment probes.

#16790 PR: Vectorize 3 loops in OptimizedScalarQuantizer
7.8 hours ago  0 comments  0 votes  0 watches  github-actions[bot], slow-J
Starting as a draft, as I want to make sure this approach is valid first. I was looking into HNSW indexing performance (also https://github.com/apache/lucene/pull/16631 ) ... OptimizedScalarQuantizer.scalarQuantize centers the vector, then runs a coordinate descent to pick ...

#12892: Remove all deprecated IndexSearcher#search(Query, Collector) usage / methods in the next major ...
8.9 hours ago  7 comments  0 votes  0 watches  gaobinlong, github-project-automation[bot], gsmiller, javanna, msfroh, msokolov, romseygeek, sgup432, vijaykriishna, zacharymorn
Description As a follow-up of https://github.com/apache/lucene/issues/11041, we would like to ... A list of the leftover usages follows:- [x] facet: FacetsCollector (ongoing discussion at #13725, ...
    msokolov 8.9 hours ago:  I went a bit deeper and started writing code to aggregate per-term scores across multiple ... However, one wrinkle is that AVG collection is supported, but this requires deep changes since it ...
    msokolov 1.1 days ago:  I made a stab at converting the JoinUtils over to CollectionManager, but it is not trivial as I ... I wonder if we could somehow disable the multithreading. Maybe the CollectorManager could signal its need to run on a single thread somehow?

[35.0 msec search, 36.7 msec total]