Fix loc len #45

simon-mo · 2018-07-17T21:48:10Z

What do these changes do?

Fix iloc issue. The core fix is the line:

row_lengths_oid = ray.put(np.array(row_lengths))

and other similar lines. IndexMetadata accepts numpy array instead of list.

Related issue number

#43

passes git diff upstream/master -u -- "*.py" | flake8 --diff
Add to test suite

kunalgosar · 2018-07-17T21:57:54Z

modin/pandas/indexing.py

        row_metadata_view = _IndexMetadata(
-            coord_df_oid=row_lookup, lengths_oid=row_lengths)
+            coord_df_oid=row_lookup, lengths_oid=row_lengths_oid)


Can we modify the metadata class to take in an np.array here, as well as an OID, instead of using ray.put?

I think it should be just oid here because the keyword arg is _oid. When we refactor IndexMetaData (which will be another PR and coming soon), we can make it taking np.array. I don't want this to be a refactoring PR.

Can we file an issue or mark a todo in the code in that case?

Yep. I think @devin-petersohn will soon write up a design doc for the new index (using two B+Tree) that will basically re-write _IndexMetaData.

Yes, refactoring index is WIP. Exact datastructure is still an open discussion 😉

kunalgosar · 2018-07-17T21:58:15Z

modin/pandas/indexing.py


                col_item_index += col_len
            row_item_index += row_len

+        if self.is_view:
+            warn(_SETTING_WITHOUT_COPYING_WARING)


Can this warning be imported from pandas?

Sadly, no... Pandas uses inline warning text:
https://github.com/pandas-dev/pandas/blob/27ebb3e1e40513ad5f8919a5bbc7298e2e070a39/pandas/core/generic.py#L2778-L2805

kunalgosar · 2018-07-17T22:01:20Z

modin/pandas/dataframe.py

@@ -123,7 +125,10 @@ def __init__(self, data=None, index=None, columns=None, dtype=None,
            if block_partitions is not None:
                axis = 0
                # put in numpy array here to make accesses easier since it's 2D
-                self._block_partitions = np.array(block_partitions)
+                if not isinstance(block_partitions, np.ndarray):


Can this check be done in _fix_blocks_dimensions?

simon-mo · 2018-07-20T07:22:27Z

Closed in favor of #55

…ze_count Izamyati/groupby.size()/count()

…a service Signed-off-by: Devin Petersohn <[email protected]> Fixes to pass CI + docs for io.py Update implementation Signed-off-by: Devin Petersohn <[email protected]> Fix some things Signed-off-by: Devin Petersohn <[email protected]> Lint fixes Fix put Signed-off-by: Devin Petersohn <[email protected]> Clean up and add new details Signed-off-by: Devin Petersohn <[email protected]> Use fsspec to get full path and allow URLs Signed-off-by: Devin Petersohn <[email protected]> Add lazy loc Signed-off-by: Devin Petersohn <[email protected]> fixes for tests porting more tests more fixes moar fixes Raise exception Signed-off-by: Devin Petersohn <[email protected]> Lint fixes Return Python as the default modin engine Handle indexing case for client qc Call fast path for __getitem__ if not lazy Remove user warning for Python-engine fall back Add init Signed-off-by: Devin Petersohn <[email protected]> Implement free as a no-op Signed-off-by: Devin Petersohn <[email protected]> Add support for replace - client side Fix a couple of issues with Client Signed-off-by: Devin Petersohn <[email protected]> Throw errors on to_pandas Signed-off-by: Devin Petersohn <[email protected]> Do not default to pandas for str_repeat Add support for 18 datetime functions/properties Fix columns caching when renaming columns Fix test_query: put backticks back for col names Add support for astype -- client side hard coded changes for functions Client support for str_(en/de)code, to_datetime Add all missing query compiler methods. Signed-off-by: mvashishtha <[email protected]> Fix getitem_column_array and take_2d. Signed-off-by: mvashishtha <[email protected]> Fix getitem_column_array and take_2d. Signed-off-by: mvashishtha <[email protected]> Fix again. Signed-off-by: mvashishtha <[email protected]> Fix more bugs. Signed-off-by: mvashishtha <[email protected]> More fixes. Signed-off-by: mvashishtha <[email protected]> Fix more bugs-- pushdown tests test_dates and test_pivot still broken due to service bugs. Signed-off-by: mvashishtha <[email protected]> Fix typo. Note drop() broken because service requires you to specify both argument and client QC at base of this PR uses default Nones. Signed-off-by: mvashishtha <[email protected]> Add query compiler class. Signed-off-by: mvashishtha <[email protected]> Testing a commit Initial changes for adding support for Expanding FEAT Support for rolling.sem FEAT support for Expanding sum, min, max, mean, var, std, count, sem Removing extratenous comment REFACTOR: Remove defaults to pandas at API layer and add some corresponding client QC methods. Signed-off-by: mvashishtha <[email protected]> Add more methods. Signed-off-by: mvashishtha <[email protected]> Fix expanding. Signed-off-by: mvashishtha <[email protected]> Add ewm. Signed-off-by: mvashishtha <[email protected]> Revert whitespace. Signed-off-by: mvashishtha <[email protected]> Fix to_numpy by making it like to_pandas. Signed-off-by: mvashishtha <[email protected]> Remove extra to_numpy. Signed-off-by: mvashishtha <[email protected]> Pass kwargs Signed-off-by: mvashishtha <[email protected]> Fix DataFrame import for isin. Signed-off-by: mvashishtha <[email protected]> Fix again. Signed-off-by: mvashishtha <[email protected]> Remove breakpoint Signed-off-by: mvashishtha <[email protected]> Tell if series. Signed-off-by: mvashishtha <[email protected]> Fix client qc. Signed-off-by: mvashishtha <[email protected]> Add self_is_series. Signed-off-by: mvashishtha <[email protected]> FIX: Set numeric_only to True in groupby quantile Add some comments Fix str_cat/fullmatch/removeprefix/removesuffix/translate/wrap (modin-project#44) * Fix str_cat/fullmatch/removeprefix/removesuffix/translate/wrap * Update modin/core/storage_formats/base/query_compiler.py Co-authored-by: Mahesh Vashishtha <[email protected]> * Update modin/pandas/series_utils.py Co-authored-by: Mahesh Vashishtha <[email protected]> * Update modin/core/storage_formats/base/query_compiler.py Co-authored-by: Mahesh Vashishtha <[email protected]> Co-authored-by: Mahesh Vashishtha <[email protected]> FEAT Support expanding.aggregate (modin-project#45) Fix at_time and between_time. (modin-project#43) Signed-off-by: mvashishtha <[email protected]> Signed-off-by: mvashishtha <[email protected]> Add QC method for groupby.sem (modin-project#47) * FEAT: Add partial support for groupby.sem() * Add sem changes to groupby Fix nlargest and nsmallest Series support (modin-project#46) * Fix nlargest and smallest support Signed-off-by: Naren Krishna <[email protected]> Remove client query compiler's columnarize. (modin-project#48) Signed-off-by: mvashishtha <[email protected]> Signed-off-by: mvashishtha <[email protected]> Fix info and set memory_usage=False. (modin-project#49) Signed-off-by: mvashishtha <[email protected]> Signed-off-by: mvashishtha <[email protected]> POND-815 fixes for 21 column dataset (modin-project#50) * POND-815 fixes for 21 column dataset * Update modin/pandas/base.py Co-authored-by: helmeleegy <[email protected]> --------- Co-authored-by: helmeleegy <[email protected]> Bring in upstream series binary operation fix 6d5545f… (modin-project#52) * Bring in upstream series binary operation fix 6d5545f. Signed-off-by: mvashishtha <[email protected]> * Update modin/pandas/series.py Co-authored-by: Karthik Velayutham <[email protected]> --------- Signed-off-by: mvashishtha <[email protected]> Co-authored-by: Karthik Velayutham <[email protected]> Support groupby first/last (modin-project#53) Signed-off-by: Naren Krishna <[email protected]> FEAT: Add initial partial support for groupby.cumcount() (modin-project#54) * FEAT: Add partial support for cumcount * Remove the set_index_name * Squeeze the result * Write cumcount name to None * Can't set dtype to int64 Fix resample sum, prod, size (modin-project#56) Signed-off-by: Naren Krishna <[email protected]> POND-184: fix describe and simplify query compiler interface (modin-project#55) * Fix describe Signed-off-by: mvashishtha <[email protected]> * Pass datetime_is_numeric. Signed-off-by: mvashishtha <[email protected]> --------- Signed-off-by: mvashishtha <[email protected]> Fix dt_day_of_week/day_of_year, str_cat/extract/partition/replace/rpartition (modin-project#51) * Fix dt_day_of_week/day_of_year, str_partition/replace/rpartition * Fix str_extract Revert "Fix dt_day_of_week/day_of_year, str_cat/extract/partition/replace/rpartition (modin-project#51)" (modin-project#58) This reverts commit f7a31ab. Revert "Revert "Fix dt_day_of_week/day_of_year, str_cat/extract/partition/replace/rpartition (modin-project#51)" (modin-project#58)" (modin-project#60) This reverts commit ad9231d. Add query compiler method for groupby.prod() (modin-project#57) Signed-off-by: Naren Krishna <[email protected]> FEAT: Add support for groupby.head and groupby.tail (modin-project#61) * FEAT: Add support for groupby.head and groupby.tail * Change _change_index FEAT: Add partial support for groupby.nth (modin-project#62) FIX: Push first and last down to query compiler. (modin-project#64) * FIX: Push first and last down to query compiler. Signed-off-by: mvashishtha <[email protected]> * Fix last. Signed-off-by: mvashishtha <[email protected]> --------- Signed-off-by: mvashishtha <[email protected]> FEAT: Add partial support for groupby.ngroup (modin-project#65) * FEAT: Add partial support for groupby.ngroup * Name of result should be none for now Add client support for SeriesGroupby unique, nsmallest, nlargest (modin-project#63) * Add client support for SeriesGroupby unique, nsmallest, nlargest Signed-off-by: Naren Krishna <[email protected]> --------- Signed-off-by: Naren Krishna <[email protected]> Push memory_usage entirely to query compiler [change is not to be upstreamed to Modin] (modin-project#66) * Fix dataframe memory usage. Signed-off-by: mvashishtha <[email protected]> * Fix series memory_usage() the same way. Signed-off-by: mvashishtha <[email protected]> --------- Signed-off-by: mvashishtha <[email protected]> FIX: allow updating backend query compilers in place. (modin-project#67) * FIX: Mutate client query compiler columns and index in the service. Motivation: Align axis update semantics across query compilers. In the base query compiler and even our service's query compiler, you can update the index and columns in place. However, the service gives no way to update axes of a query compiler. Right now, for inplace updates, service exposes an extra method rename(), and client query compiler uses this to get the id of a new compiler with updated axis, and then updates its id ID of the new query compiler. This change might be the first to make the service present a mutable interface for a backend query compiler. That seems safe to me, except I had to make copy() get a new query compiler copied from the old query compiler, because we can't let updates to the new query compiler change the original (or vice versa). Signed-off-by: mvashishtha <[email protected]> * Add a comment. Signed-off-by: mvashishtha <[email protected]> --------- Signed-off-by: mvashishtha <[email protected]> FEAT replace groupby.fillna with a simpler logic (modin-project#68) * FEAT Support expanding.aggregate * Replaced groupby.fillna logic with a simpler one * Fix in groupby.fillna. Work object was causing problems. * Only need to change _check_index_name to _check_index * Removed commented out code.

simon-mo added 5 commits July 17, 2018 13:15

Fix view -> origin propagation

d5db728

Add warning to df_view

8fe1389

Remove mask_block_partitions

2ee2ab2

Fix the last IndexMetadata constructor

4275cdc

Flake8

5162f8d

kunalgosar requested changes Jul 17, 2018

View reviewed changes

simon-mo mentioned this pull request Jul 20, 2018

[Indexing] Fix #43 & Copy DataFrame for view #55

Merged

1 task

simon-mo closed this Jul 20, 2018

dchigarev pushed a commit to dchigarev/modin that referenced this pull request Aug 25, 2020

Merge pull request modin-project#45 from intel-go/izamyati/groupby_si…

236c2c6

…ze_count Izamyati/groupby.size()/count()

mdatre added a commit to mdatre/modin that referenced this pull request Jan 21, 2023

FEAT Support expanding.aggregate (modin-project#45)

b074fa3

vnlitvinov pushed a commit to vnlitvinov/modin that referenced this pull request Feb 13, 2023

FEAT Support expanding.aggregate (modin-project#45)

f369b4d

vnlitvinov pushed a commit to vnlitvinov/modin that referenced this pull request Mar 16, 2023

FEAT Support expanding.aggregate (modin-project#45)

5a11ba9

vnlitvinov pushed a commit to vnlitvinov/modin that referenced this pull request Mar 16, 2023

FEAT Support expanding.aggregate (modin-project#45)

04b5379

vnlitvinov pushed a commit to vnlitvinov/modin that referenced this pull request Mar 16, 2023

FEAT Support expanding.aggregate (modin-project#45)

0256161

noloerino pushed a commit to noloerino/modin that referenced this pull request May 2, 2023

FEAT Support expanding.aggregate (modin-project#45)

d61a1db

noloerino pushed a commit to noloerino/modin that referenced this pull request May 3, 2023

FEAT Support expanding.aggregate (modin-project#45)

e16a7b8

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Fix loc len #45

Fix loc len #45

simon-mo commented Jul 17, 2018

kunalgosar Jul 17, 2018

simon-mo Jul 19, 2018

kunalgosar Jul 19, 2018

simon-mo Jul 19, 2018

pschafhalter Jul 19, 2018 •

edited

Loading

kunalgosar Jul 17, 2018

simon-mo Jul 19, 2018

kunalgosar Jul 19, 2018

kunalgosar Jul 17, 2018

simon-mo Jul 19, 2018

simon-mo commented Jul 20, 2018

Fix loc len #45

Fix loc len #45

Conversation

simon-mo commented Jul 17, 2018

What do these changes do?

Related issue number

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

pschafhalter Jul 19, 2018 • edited Loading

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

simon-mo commented Jul 20, 2018

pschafhalter Jul 19, 2018 •

edited

Loading